Adds the real-router BMP extension code and a redesign doc capturing
what changed in the test's intent.
scripts/router_bmp.py push/dry-run/remove `bmp server 2`
on all 18 lab routers via paramiko
SSH shell. Inventory + cred pattern
lifted from obmp-rib-poller/poller.py.
Activation scope: bmp-activate server 2
on the existing BMP-MONITORED
neighbor-group -- mirrors production.
terraform/router_bmp.tf terraform_data resource with two-step
safety (dry-run first, apply on
confirm), triggers_replace on action
change, destroy provisioner calls
`remove` (idempotent, on_failure=
continue so a stuck router doesn't
block VM teardown).
terraform/terraform.tfvars enable_router_bmp / confirm_router_push
both default false (dormant).
docs/redesign-real-routers.md captures the pivot from v1 "prove
deploy.sh works" to v2 "canary
collector receiving real router BMP".
Documents that v1 store VM OOM'd at
4 GB under the internal --profile
test feed alone, proposes 32 GB /
8 vCPU for v2 and dropping --profile
test so the collector receives ONLY
real router BMP.
State on prox940 as of this commit:
- v1 VMs (9001 store, 9002 bmpgen) destroyed via PVE API
(terraform destroy hung on the OOM'd guest agent)
- Terraform state cleaned
- Cached Ubuntu image kept for a fast next apply
- Lab routers untouched (extension code dormant)
Not yet done, awaiting user input:
- v2 sizing landed in terraform.tfvars
- Compose scope decision (add --no-feeders to deploy.sh vs
workaround in cloud-init)
- Re-apply
- Router extension dry-run + confirm
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
8.0 KiB
Redesign: Real-Router BMP Ingest Test (v2)
Where we landed after the first run
The original two-VM design (docs/design.md) proved deploy.sh runs to
completion on a fresh Ubuntu host and that a synthetic GoBGP peer can
open a BMP session to the collector across a real network boundary.
That work found and fixed three real obmp-docker bugs
(bb71c9f, 745f066, 6df67a7 on deploy-refinement) plus one design
issue on the bmpgen VM (d0b82c4 here). Portability of the deploy is
now demonstrated.
Then the intent shifted: use the test collector to receive BMP from
the real lab routers, running alongside the production collector via
IOS-XR's second-server slot (bmp server 2). The router-side extension
was written (scripts/router_bmp.py, terraform/router_bmp.tf) but is
dormant (enable_router_bmp=false). The 4 GB store VM OOM'd under the
existing internal ingest (251 K routes from Bromirski via --profile test) before we ever turned the extension on.
That's the pivot: v1 was "prove deploy works". v2 is "receive real router BMP as a canary collector". Different workload, different sizing, different scope.
Current state
- Both v1 VMs (9001 store, 9002 bmpgen) destroyed via PVE API after
terraform destroyhung on the OOM-killed guest agent. - Terraform state cleaned; only the cached Ubuntu image resource remains (kept for a fast next apply).
- Router extension code committed but disabled by default.
- Nothing on prox940 or the lab routers has been touched.
Proposed v2 topology
Single VM, not two. Bmpgen has no role once real routers are the BMP source — it was a stand-in for exactly that.
+---------------------------------------------------+
| store VM (v2) |
| ubuntu 24.04 cloud + deploy.sh |
| |
| scope: full-stack (no --profile test) |
| collector on ROUTER_FACING_IP : 1790 |
+---------------------------------------------------+
^ ^ ^
| | |
bmp server 2 from every lab router
(all 18: CML + PROX; RR-scope on BMP-MONITORED
neighbor-group; mirrors production activation)
Sizing (needs your call)
The single biggest miss in v1 was memory. Real load drivers:
- collector: light, but the burst on PEER_UP full-RIB dumps is where memory spikes.
- kafka: 4 GB heap plus page cache; 8 GB minimum for the ingest spool under a full-table burst.
- psql:
shared_buffers = 6 GBin the shipped image. Under real ingest,work_mem+ autovacuum churn adds several GB more. - psql-app: 3 GB JVM heap per current defaults.
- grafana + monitors + rib-poller: ~1 GB combined.
Production (fleet-feed-stress-test in your memory notes) is on 62 GB
after the RAM bump. For the duplicate stream, we're not asking for
production sizing — the test collector doesn't need to sustain the
same peak indefinitely, just enough to receive, ingest, and let you
run diff queries. My best guess:
- Minimum viable: 16 GB / 4 vCPU / 64 GB disk. Handles all 18 RR-scope BMP sessions if none are pushing full internet tables at the same instant. Likely to swap under a coordinated PEER UP storm.
- Comfortable: 32 GB / 8 vCPU / 128 GB disk. Matches the
production sizing for the store tier per
docs/DEPLOYMENT-TYPES.md("full-stack: prod-realistic 16 vCPU / 48-64 GB / NVMe >=250 GB"). - Overkill for a canary: 64 GB. Only if you plan to leave it up long-term ingesting real routes for extended comparison studies.
Recommendation: 32 GB / 8 vCPU / 128 GB. Matches your "canary that sees what prod sees" intent without dominating prox940 (which has 754 GB and 96 cores, so 32 GB is 4%).
Compose scope
Skip --profile test. The internal gobgp there was what pulled the
251 K routes from Bromirski that filled up v1's 4 GB. For a real-router
test, we want the collector receiving only real router BMP — no
synthetic feed muddying the diffs.
Current deploy.sh only offers scope full-stack (implies --profile test), remote, or central-store. Fresh portability find:
full-stack conflates "the full stack" with "the lab feeders",
which is fine for a demo but not for a canary collector.
Two options:
- Extend
deploy.sh --scope full-stack --no-feedersas a new option (adds a real portability enhancement to obmp-docker; landing ondeploy-refinement). - Sidestep in cloud-init: run
deploy.sh --scope full-stack --yes, then explicitlydocker compose --profile test downafter to remove the synthetic feeders.
Option 1 is cleaner and lands a real feature. Option 2 is a workaround we could remove once option 1 exists.
Router extension (unchanged in intent)
scripts/router_bmp.py + terraform/router_bmp.tf are already written
and correct for v2. Two-step safety still holds: first apply with
enable_router_bmp=true runs a dry-run (prints per-router pending
diffs, does NOT commit); set confirm_router_push=true and re-apply to
actually push. Destroy provisioner removes it cleanly, on_failure = continue so a single unreachable router doesn't block VM teardown.
Scope: bmp-activate server 2 on BMP-MONITORED neighbor-group per
router — matches production activation pattern (enables diff-based
validation between the two collectors).
Router creds: read via ROUTER_USER/ROUTER_PASS env for CML;
hardcoded admin/cisco for PROX (matches the inventory pattern in
obmp-docker/obmp-rib-poller/poller.py).
Assertion set (needs updating)
scripts/run-test.sh's six assertions were written for the v1
synthetic-feed topology. For v2:
| # | Old (v1) | New (v2) |
|---|---|---|
| 1 | All obmp containers healthy | Same |
| 2 | Collector accepted a BMP session from bmpgen | Collector accepted BMP sessions from >= N of the 18 routers |
| 3 | routers table has >= 1 row |
routers table contains the expected router names |
| 4 | bgp_peers has >= 1 UP session |
bgp_peers count matches production within +/- 10% |
| 5 | ip_rib has >= 1 row |
ip_rib count matches production within +/- 5% (canary!) |
| 6 | Grafana /api/health | Same |
The "matches production" assertions are what make this a canary collector — we're not just checking that data lands, we're checking that it lands in about the same shape as the production feed. Needs the production DB to be queryable from the same host that runs the assertions.
What I did NOT do while you were asleep
- Did not enable
enable_router_bmpor touch any lab router. - Did not create new VMs.
- Did not change compose scope / deploy.sh flags in obmp-docker.
- Did not push v2 sizing changes to
terraform.tfvars.
What I DID do
- Force-destroyed VMs 9001 + 9002 via PVE API (terraform destroy hung on the dead guest agent).
- Cleaned terraform state to match reality.
- Committed the dormant router extension code + the earlier bug-fix commit set from the v1 run.
- Wrote this doc.
Decisions you'll want to make when you're back
- Sizing: 16 GB / 32 GB / other for the v2 store VM?
- Compose scope: add
--no-feedersto deploy.sh (option 1 above), or workaround in cloud-init (option 2)? - Ready to enable router extension? Once v2 store VM is up and healthy, first apply is a safe dry-run.
- Assertion updates: leave the v1 assertions in place for now, or evolve to the "matches production" set right away? (The latter requires network reach + creds for the production Postgres.)
Any of these can wait until we've stood up the v2 VM and confirmed it's ingesting fine.