rp-watch: Rocket Pool Node Monitoring & Alerting Daemon
What is the work being proposed?
A lightweight, open-source Rust daemon (rp-watch) that monitors a Rocket Pool node operator’s smartnode stack in real time and pushes alerts before problems become missed duties or financial loss. It tracks:
- Minipool status changes
- Execution/consensus client sync health and downtime
- Missed or upcoming attestations
- RPL collateral ratio, with configurable warning thresholds
- Rewards period timing and claim windows
Alerts are pushed via Discord, Telegram, or generic webhook, configurable per operator. The tool runs alongside an existing smartnode installation with minimal resource overhead.
The daemon also includes an AI agent layer, built on my published crate agenticrs (a reliability/observability layer for LLM tool-calling ; retries, circuit breaking, JSON schema validation, OTel tracing), running a small, quantized model locally (via Rust-native inference bindings) rather than calling an external API. This keeps the tool self-contained, avoids recurring API costs or dependency on third-party uptime for a reliability-critical tool, and fits well alongside a node operator’s existing local-first setup. The agent periodically reasons over the collected telemetry rather than relying on static thresholds alone:
- Anomaly reasoning — the agent is given tool-call access to recent sync timing, attestation history, and collateral trend data, and periodically evaluates whether current behavior looks abnormal relative to the node’s own recent history, flagging issues before they cross hard, static thresholds.
- Natural-language alert summaries — when an alert fires, the agent drafts a plain-English diagnosis of what’s happening and a suggested next action, rather than a raw metric dump, making alerts usable by less technical operators.
- Predictive collateral commentary — given current RPL ratio trend and recent rETH price movement, the agent produces a forecast-style warning (“at this rate, you’ll likely cross your safety threshold in ~X days”) rather than a purely reactive alert.
Using a small, locally-run model with tool access (instead of training bespoke large ML models or depending on an external API) keeps this feasible within the grant window, avoids ongoing inference costs or third-party dependency for operators, and still delivers meaningfully smarter alerting than static rules alone. It also directly builds on agenticrs’s existing retry/validation/tracing infrastructure to keep the agent’s outputs reliable.
This is the first of a planned series of node-operator tooling grants; two related tools : a circuit-breaker-based client fallback tool (rp-guard) and a rETH accounting/tax export CLI (rp-ledger) — are intentionally scoped out of this round and will be proposed separately in future rounds once rp-watch has shipped and been used in production.
Is there any related work this builds off of?
Yes. The daemon architecture builds directly on my published crate daemon_rs (long-running daemon process management, already in production use), the alerting/threshold logic reuses patterns from my circuit_breaker crate, and the AI agent layer builds on my published crate agenticrs (retries, circuit breaking, JSON schema validation, and OTel tracing for LLM tool-calls). This is a purpose-built adaptation of proven, existing designs rather than work starting from zero.
Will the results of this project be entirely open source?
Yes, MIT license, matching my existing published crates.
Benefit
| Group | Benefits |
|---|---|
| Potential rETH holders | N/A — this tool is operator-facing, not holder-facing. |
| rETH holders | Improved network-wide operator reliability (fewer missed duties) supports more consistent rETH rewards accrual. |
| Potential NOs | Lowers the barrier to confidently running a node by removing the need to manually watch dashboards for problems, and by turning raw metrics into plain-English guidance — makes “operate a node” more attractive to less technical newcomers. |
| NOs | Directly reduces missed duties, downtime, and under-collateralization risk. Agent-driven anomaly reasoning and predictive collateral commentary running fully locally, with no external API dependency or recurring cost give earlier notice than static threshold alerts alone, protecting rewards and reducing slashing exposure. |
| Community | Adds to the ecosystem of open-source operator tooling; reduces reliance on operators individually building ad hoc monitoring scripts. |
| RPL holders | Reduced missed duties and under-collateralization across the operator set improves overall protocol health and reliability. |
Which other non-RPL protocols, DAOs, projects, or individuals, would stand to benefit from this grant?
The daemon architecture and alerting patterns are generalizable to any Ethereum validator or node operator setup with similar monitoring needs (other LSD protocols, solo stakers).
Work
Who is doing the work?
Mahmud Bello (GitHub: mahmudsudo), a Lagos-based systems engineer with 6+ years of production Rust and C++ experience.
What is the background of the person(s) doing the work?
6+ years of production experience in distributed systems, protocol-level infrastructure, cryptography, and developer tooling. Former CTO at StreamLivr (rebuilt failing async architecture, scaled to 5,000+ users), DevRel roles at ICP Hub Sahara West Africa and Microsandbox. Published Rust crates: daemon_rs, circuit_breaker, rate_rs, cargomon, and agenticrs (reliability and observability tooling for LLM/agent tool-calls ; retries, circuit breaking, JSON schema validation, OTel tracing). Open-source contributions to Paradigm’s Solar compiler, Ithacaxyz’s Odyssey execution layer, ingonyama-zk’s Blaze, and Pluto Labs’ Ronkathon.
What is the breakdown of the proposed work, in terms of milestones and/or deadlines?
- Week 1: Core daemon architecture — smartnode stack polling, minipool status tracking, client sync health checks.
- Week 2: Collateral ratio tracking, rewards period timing, threshold configuration system.
- Week 3: AI agent layer setup on
agenticrs— local model integration, tool-call access to telemetry, retry/validation/tracing wiring. - Week 4: Anomaly reasoning and predictive collateral commentary logic, initial prompt/response validation against historical node data.
- Week 5: Natural-language alert summary generation, alerting integrations (Discord, Telegram, generic webhook), end-to-end testing on a live/testnet node.
- Week 6: Documentation, packaging for easy install, public release, bug fixes from early tester feedback.
How is the work being tested? Is testing included in the schedule?
Yes , unit tests throughout, validation of the AI agent’s reasoning against historical node data in weeks 3-4, and integration testing against a live or testnet Rocket Pool node in weeks 5-6 to validate real alerting behavior before public release.
How will the work be maintained after delivery?
Published on GitHub under my existing account (mahmudsudo) alongside my other maintained crates. I will address community-reported issues and accept contributions post-delivery.
Costs
What is the acceptance criteria?
A working, tested, documented release of rp-watch (published to crates.io and GitHub), verified against at least one live or testnet Rocket Pool node, with alerting confirmed functional across at least one notification channel, and the AI agent’s anomaly reasoning, predictive collateral commentary, and natural-language summaries confirmed operational against real or historical node data.
What is the proposed payment schedule for the grant? How much USD $ and over what period of time is the applicant requesting?
Total ask: $15,000 USD, paid as six milestone-based installments of $2,500 each, released upon delivery of each weekly milestone (weeks 1-6 as outlined above). This covers 6 weeks of solo, focused development including local model integration and validation work. Running a small model locally (rather than training bespoke ML models or depending on a paid external API) keeps this scope feasible within the window, avoids recurring inference costs for operators, and still delivers meaningfully smarter alerting than static rules alone.
Who will directly receive the payment?
0xa8B78566b11e865C290410420c1f30C4982077CF
How will the GMC verify that the work delivered matches the proposed cadence?
Each weekly milestone delivered as a public GitHub commit/release with documentation, allowing real-time progress verification.
What alternatives or options have been considered in order to save costs for the proposed project?
Scoping to a single, narrow tool (rather than the full planned toolkit) keeps this round’s ask small and feasible for solo delivery in 4 weeks. Related tools are deferred to future rounds rather than bundled in, keeping this proposal’s cost proportionate to its scope.
Have you already been compensated by the RP protocol in any way for this work?
No.
Conflict of Interest
Does the person or persons proposing the grant have any conflicts of interest to disclose?
No.
Will the recipient of the grant, or any protocol or project in which the recipient has a vested interest (other than Rocket Pool), benefit financially if the grant is successful?
No.