obmp-portability-test/docs/redesign-real-routers.md
sam 6fc92541e9 router extension (dormant) + v2 redesign doc after OOM'd first run
Adds the real-router BMP extension code and a redesign doc capturing
what changed in the test's intent.

  scripts/router_bmp.py           push/dry-run/remove `bmp server 2`
                                  on all 18 lab routers via paramiko
                                  SSH shell. Inventory + cred pattern
                                  lifted from obmp-rib-poller/poller.py.
                                  Activation scope: bmp-activate server 2
                                  on the existing BMP-MONITORED
                                  neighbor-group -- mirrors production.

  terraform/router_bmp.tf         terraform_data resource with two-step
                                  safety (dry-run first, apply on
                                  confirm), triggers_replace on action
                                  change, destroy provisioner calls
                                  `remove` (idempotent, on_failure=
                                  continue so a stuck router doesn't
                                  block VM teardown).

  terraform/terraform.tfvars      enable_router_bmp / confirm_router_push
                                  both default false (dormant).

  docs/redesign-real-routers.md   captures the pivot from v1 "prove
                                  deploy.sh works" to v2 "canary
                                  collector receiving real router BMP".
                                  Documents that v1 store VM OOM'd at
                                  4 GB under the internal --profile
                                  test feed alone, proposes 32 GB /
                                  8 vCPU for v2 and dropping --profile
                                  test so the collector receives ONLY
                                  real router BMP.

State on prox940 as of this commit:
  - v1 VMs (9001 store, 9002 bmpgen) destroyed via PVE API
    (terraform destroy hung on the OOM'd guest agent)
  - Terraform state cleaned
  - Cached Ubuntu image kept for a fast next apply
  - Lab routers untouched (extension code dormant)

Not yet done, awaiting user input:
  - v2 sizing landed in terraform.tfvars
  - Compose scope decision (add --no-feeders to deploy.sh vs
    workaround in cloud-init)
  - Re-apply
  - Router extension dry-run + confirm

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-21 07:21:03 -07:00

8.0 KiB

Redesign: Real-Router BMP Ingest Test (v2)

Where we landed after the first run

The original two-VM design (docs/design.md) proved deploy.sh runs to completion on a fresh Ubuntu host and that a synthetic GoBGP peer can open a BMP session to the collector across a real network boundary. That work found and fixed three real obmp-docker bugs (bb71c9f, 745f066, 6df67a7 on deploy-refinement) plus one design issue on the bmpgen VM (d0b82c4 here). Portability of the deploy is now demonstrated.

Then the intent shifted: use the test collector to receive BMP from the real lab routers, running alongside the production collector via IOS-XR's second-server slot (bmp server 2). The router-side extension was written (scripts/router_bmp.py, terraform/router_bmp.tf) but is dormant (enable_router_bmp=false). The 4 GB store VM OOM'd under the existing internal ingest (251 K routes from Bromirski via --profile test) before we ever turned the extension on.

That's the pivot: v1 was "prove deploy works". v2 is "receive real router BMP as a canary collector". Different workload, different sizing, different scope.

Current state

  • Both v1 VMs (9001 store, 9002 bmpgen) destroyed via PVE API after terraform destroy hung on the OOM-killed guest agent.
  • Terraform state cleaned; only the cached Ubuntu image resource remains (kept for a fast next apply).
  • Router extension code committed but disabled by default.
  • Nothing on prox940 or the lab routers has been touched.

Proposed v2 topology

Single VM, not two. Bmpgen has no role once real routers are the BMP source — it was a stand-in for exactly that.

   +---------------------------------------------------+
   |             store VM (v2)                         |
   |   ubuntu 24.04 cloud + deploy.sh                  |
   |                                                   |
   |   scope: full-stack (no --profile test)           |
   |   collector on ROUTER_FACING_IP : 1790            |
   +---------------------------------------------------+
              ^         ^         ^
              |         |         |
         bmp server 2 from every lab router
         (all 18: CML + PROX; RR-scope on BMP-MONITORED
         neighbor-group; mirrors production activation)

Sizing (needs your call)

The single biggest miss in v1 was memory. Real load drivers:

  • collector: light, but the burst on PEER_UP full-RIB dumps is where memory spikes.
  • kafka: 4 GB heap plus page cache; 8 GB minimum for the ingest spool under a full-table burst.
  • psql: shared_buffers = 6 GB in the shipped image. Under real ingest, work_mem + autovacuum churn adds several GB more.
  • psql-app: 3 GB JVM heap per current defaults.
  • grafana + monitors + rib-poller: ~1 GB combined.

Production (fleet-feed-stress-test in your memory notes) is on 62 GB after the RAM bump. For the duplicate stream, we're not asking for production sizing — the test collector doesn't need to sustain the same peak indefinitely, just enough to receive, ingest, and let you run diff queries. My best guess:

  • Minimum viable: 16 GB / 4 vCPU / 64 GB disk. Handles all 18 RR-scope BMP sessions if none are pushing full internet tables at the same instant. Likely to swap under a coordinated PEER UP storm.
  • Comfortable: 32 GB / 8 vCPU / 128 GB disk. Matches the production sizing for the store tier per docs/DEPLOYMENT-TYPES.md ("full-stack: prod-realistic 16 vCPU / 48-64 GB / NVMe >=250 GB").
  • Overkill for a canary: 64 GB. Only if you plan to leave it up long-term ingesting real routes for extended comparison studies.

Recommendation: 32 GB / 8 vCPU / 128 GB. Matches your "canary that sees what prod sees" intent without dominating prox940 (which has 754 GB and 96 cores, so 32 GB is 4%).

Compose scope

Skip --profile test. The internal gobgp there was what pulled the 251 K routes from Bromirski that filled up v1's 4 GB. For a real-router test, we want the collector receiving only real router BMP — no synthetic feed muddying the diffs.

Current deploy.sh only offers scope full-stack (implies --profile test), remote, or central-store. Fresh portability find: full-stack conflates "the full stack" with "the lab feeders", which is fine for a demo but not for a canary collector.

Two options:

  1. Extend deploy.sh --scope full-stack --no-feeders as a new option (adds a real portability enhancement to obmp-docker; landing on deploy-refinement).
  2. Sidestep in cloud-init: run deploy.sh --scope full-stack --yes, then explicitly docker compose --profile test down after to remove the synthetic feeders.

Option 1 is cleaner and lands a real feature. Option 2 is a workaround we could remove once option 1 exists.

Router extension (unchanged in intent)

scripts/router_bmp.py + terraform/router_bmp.tf are already written and correct for v2. Two-step safety still holds: first apply with enable_router_bmp=true runs a dry-run (prints per-router pending diffs, does NOT commit); set confirm_router_push=true and re-apply to actually push. Destroy provisioner removes it cleanly, on_failure = continue so a single unreachable router doesn't block VM teardown.

Scope: bmp-activate server 2 on BMP-MONITORED neighbor-group per router — matches production activation pattern (enables diff-based validation between the two collectors).

Router creds: read via ROUTER_USER/ROUTER_PASS env for CML; hardcoded admin/cisco for PROX (matches the inventory pattern in obmp-docker/obmp-rib-poller/poller.py).

Assertion set (needs updating)

scripts/run-test.sh's six assertions were written for the v1 synthetic-feed topology. For v2:

# Old (v1) New (v2)
1 All obmp containers healthy Same
2 Collector accepted a BMP session from bmpgen Collector accepted BMP sessions from >= N of the 18 routers
3 routers table has >= 1 row routers table contains the expected router names
4 bgp_peers has >= 1 UP session bgp_peers count matches production within +/- 10%
5 ip_rib has >= 1 row ip_rib count matches production within +/- 5% (canary!)
6 Grafana /api/health Same

The "matches production" assertions are what make this a canary collector — we're not just checking that data lands, we're checking that it lands in about the same shape as the production feed. Needs the production DB to be queryable from the same host that runs the assertions.

What I did NOT do while you were asleep

  • Did not enable enable_router_bmp or touch any lab router.
  • Did not create new VMs.
  • Did not change compose scope / deploy.sh flags in obmp-docker.
  • Did not push v2 sizing changes to terraform.tfvars.

What I DID do

  • Force-destroyed VMs 9001 + 9002 via PVE API (terraform destroy hung on the dead guest agent).
  • Cleaned terraform state to match reality.
  • Committed the dormant router extension code + the earlier bug-fix commit set from the v1 run.
  • Wrote this doc.

Decisions you'll want to make when you're back

  1. Sizing: 16 GB / 32 GB / other for the v2 store VM?
  2. Compose scope: add --no-feeders to deploy.sh (option 1 above), or workaround in cloud-init (option 2)?
  3. Ready to enable router extension? Once v2 store VM is up and healthy, first apply is a safe dry-run.
  4. Assertion updates: leave the v1 assertions in place for now, or evolve to the "matches production" set right away? (The latter requires network reach + creds for the production Postgres.)

Any of these can wait until we've stood up the v2 VM and confirmed it's ingesting fine.