obmp-docker/deploy.sh
Sam b480615ca2 BMP port 1790 everywhere: parameterize the published collector port (finding 12)
Decision landed: routers target 1790 (IANA BMP port); the collector keeps
listening on 5000 in-container. One .env variable now drives every hop:

- compose publishes ${ROUTER_FACING_PORT:-1790}:5000 (was hardcoded 5000:5000)
- gobgpd.conf.tmpl (+ evpn) gains __BMP_PORT__; setup.sh renders it from
  ROUTER_FACING_PORT -- gobgp is host-networked, so it must chase the
  published port or the full-table feeder silently breaks
- deploy.sh + scripts/wsl-portproxy.ps1: the WSL side now publishes the
  router-facing port itself, so the portproxy forwards the same port
  end-to-end instead of down-mapping to 5000
- cml/proxmox_bmp_config.py reads ROUTER_FACING_IP/PORT from .env (HOST_IP
  fallback for pre-key .envs); xrd-node-definition.yaml baked port -> 1790
- .env.example documents ROUTER_FACING_IP/PORT; docs updated, finding 12
  closed on the repo side (routers still on 5000 need a config re-apply)

Verified live: collector republished 0.0.0.0:1790->5000, gobgp BMP session
Up to HOST_IP:1790 with the RIB dump flowing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 16:33:52 -07:00

755 lines
36 KiB
Bash
Executable File

#!/usr/bin/env bash
#
# deploy.sh - unified interactive OpenBMP stack deploy.
#
# One front door for standing up (or rebuilding) the obmp-docker stack on:
# - a WSL2 test box,
# - a native-Linux Docker host (central "store" node), or
# - a native-Linux REMOTE COLLECTOR that ingests BMP near the routers and
# forwards to a central store's Kafka (scale-out ingestion).
#
# It reconciles the host-specific values in .env FIRST (the "works everywhere"
# part), shows you a plan, and only then runs setup.sh + brings the stack up
# in stages. Interactive by default; fully scriptable with flags + --yes.
#
# ---------------------------------------------------------------------------
# USAGE
# ./deploy.sh # interactive; auto-detect + prompts
# ./deploy.sh --wsl | --prod # force host type (still prompts rest)
# ./deploy.sh --host-ip 10.0.0.50 # preset HOST_IP (internal stack IP)
# ./deploy.sh --router-ip 10.0.0.9 --router-port 5000 # BMP-facing target
# ./deploy.sh --auth local|authelia
# ./deploy.sh --scope core|feeders|full|standalone
# ./deploy.sh --central-kafka 10.0.0.50:9092 # for --scope standalone
# ./deploy.sh --reset # wipe data tree first (guarded on prod)
# ./deploy.sh --prod --yes # unattended (needs enough presets)
# REPO_DIR=/opt/obmp-docker ./deploy.sh
#
# SCOPE / DEPLOYMENT LEVELS (what each includes, excludes, ballpark sysreqs):
#
# core Collector core: zookeeper, kafka, psql, collector, psql-app,
# grafana, whois. A complete single-node store.
# INCLUDES: BMP ingest + Kafka bus + Postgres/TimescaleDB + UI.
# EXCLUDES: lab feeders, Authelia.
# SYSREQS: DIMENSION FOR BURST, NOT STEADY STATE. Postgres write
# latency is the bottleneck that saturates the whole host - the
# collector process itself is light, but it is NOT the sizing
# driver. See "BMP burst sizing" below.
# - small/lab (few peers, partial tables): 4 vCPU, 8-16 GB,
# 80+ GB SSD.
# - FULL-TABLE lab (GoBGP full v4/v6 reflected through RRs to
# several clients - the observed hammer case): 8+ vCPU,
# 32-64 GB RAM, NVMe with real IOPS headroom (>=5000). The
# load is full-table x monitored-sessions, not peer count.
# - 50-100 peers / full tables: 8-16 vCPU, 32-64 GB RAM,
# NVMe 1-2 TB.
# Storage scales with PREFIXES: ~1 GB/peer full internet table
# (+~50 MB/day timeseries); internal peers far less.
#
# feeders core + --profile test lab feeders (exabgp, gobgp, traffic-gen,
# influxdb, telegraf, rib-poller, churn/kafka-lag monitors).
# INCLUDES: everything in core + synthetic BGP/BMP generators +
# gNMI telemetry sink for lab exercise.
# EXCLUDES: Authelia.
# SYSREQS: core + ~4-8 GB RAM headroom for feeders (exabgp full
# table can hold ~900K route objects). 6+ vCPU comfortable. NOTE:
# synthetic feeders let you REPRODUCE the burst that sizes core -
# use them to load-test before pointing real routers at it.
#
# full feeders + --profile auth (Authelia + portal). Front-door auth.
# INCLUDES: everything in feeders + SSO/reverse-proxy.
# EXCLUDES: nothing.
# REQUIRES: auth mode = authelia and a real OBMP_DOMAIN.
# SYSREQS: as feeders (+ negligible for Authelia/nginx).
#
# standalone REMOTE COLLECTOR only: zookeeper + kafka + collector locally,
# forwarding parsed/raw data to a CENTRAL store's Kafka. No local
# Postgres/psql-app/Grafana - the central node owns the store.
# INCLUDES: BMP ingest near the routers; produces to central bus.
# EXCLUDES: Postgres, psql-app, Grafana, whois, feeders, Authelia.
# WHY: the collector is CPU-light and stateless; the DB is the
# heavy tier that saturates under burst. Splitting collectors out
# near the routers spreads the connect/flap burst across nodes so
# each Kafka spool absorbs only its own routers' dumps, instead of
# 18 full tables hitting one host at once (the exact failure mode
# observed on a single-node lab). Converge all data in one central
# Kafka -> one Postgres.
# SYSREQS (per remote node): 2-4 vCPU, 4-8 GB RAM, ~20-40 GB fast
# disk for the Kafka spool (must buffer a full connect/flap burst
# from its local routers without dropping). The CENTRAL store is
# still sized per the core burst guidance above.
# REQUIRES: --central-kafka HOST:PORT (the store node's Kafka).
#
# ---------------------------------------------------------------------------
# BMP BURST SIZING (why steady-state numbers mislead)
# BMP is event-driven and bursty. Steady state is trivial; the sizing events
# are (1) initial RIB dump on PEER_UP, (2) all routers reconnecting together
# (collector restart / RR failover), (3) reconvergence/flap churn.
#
# THE MULTIPLIER THAT ACTUALLY BITES - full table x sessions:
# A full v4+v6 table is ~960k v4 + ~220k v6 ~= 1.18M paths (~350-500 MB in
# OpenBMP's parsed Postgres form, several GB with history/timestamps). In an
# RR topology the SAME full table is reflected to every client, so BMP
# ingests it ONCE PER MONITORED SESSION. Observed lab: GoBGP (AS 65100) pulls
# full v4/v6 and injects into an iBGP mesh; RRs (CML/PROX-CORE-01/02) reflect
# to 7 clients; BMP-activating the RR-CLIENTS group means OpenBMP absorbs the
# full table x (RRs x clients) sessions at once. That is what saturates a
# single host - not "18 x 100k". It is a handful of speakers each carrying a
# FULL table, monitored across many reflected sessions.
# Root cause when the whole host melts: Postgres write latency + WRITE
# AMPLIFICATION (every reflected copy re-persisted) backpressuring
# psql-app -> Kafka -> the BMP TCP sessions.
#
# Mitigations: NVMe with real IOPS; BMP-monitor pre-policy on the RRs only
# (not every client session) to avoid ingesting N reflected copies of the
# same table; stagger connects (initial-refresh delay/spread); raise
# KAFKA/PSQL_APP mem; or split ingest with --scope standalone collectors.
# ---------------------------------------------------------------------------
#
# HOST TYPE affects default HOST_IP source, default auth/scope, reset guarding,
# and post-deploy notes (WSL prints portproxy; prod prints router guidance).
# ---------------------------------------------------------------------------
set -euo pipefail
# --- args -------------------------------------------------------------------
REPO_DIR="${REPO_DIR:-$(cd "$(dirname "$0")" && pwd)}"
ARG_HOSTTYPE="" # wsl | prod (empty = auto-detect)
ARG_HOST_IP=""
ARG_ROUTER_IP=""
ARG_ROUTER_PORT=""
ARG_AUTH="" # local | authelia
ARG_SCOPE="" # full-stack | remote | central-store (aliases accepted)
ARG_CENTRAL_KAFKA="" # host:port for --scope standalone
ARG_RESET=0
ARG_YES=0
while [ $# -gt 0 ]; do
case "$1" in
--wsl) ARG_HOSTTYPE="wsl" ;;
--prod) ARG_HOSTTYPE="prod" ;;
--host-ip) ARG_HOST_IP="${2:?--host-ip needs a value}"; shift ;;
--host-ip=*) ARG_HOST_IP="${1#*=}" ;;
--router-ip) ARG_ROUTER_IP="${2:?--router-ip needs a value}"; shift ;;
--router-ip=*) ARG_ROUTER_IP="${1#*=}" ;;
--router-port) ARG_ROUTER_PORT="${2:?--router-port needs a value}"; shift ;;
--router-port=*) ARG_ROUTER_PORT="${1#*=}" ;;
--auth) ARG_AUTH="${2:?--auth needs a value}"; shift ;;
--auth=*) ARG_AUTH="${1#*=}" ;;
--scope) ARG_SCOPE="${2:?--scope needs a value}"; shift ;;
--scope=*) ARG_SCOPE="${1#*=}" ;;
--central-kafka) ARG_CENTRAL_KAFKA="${2:?--central-kafka needs host:port}"; shift ;;
--central-kafka=*) ARG_CENTRAL_KAFKA="${1#*=}" ;;
--reset) ARG_RESET=1 ;;
--yes|-y) ARG_YES=1 ;;
-h|--help) grep '^#' "$0" | sed 's/^#//'; exit 0 ;;
*) echo "Unknown arg: $1" >&2; exit 1 ;;
esac
shift
done
log() { printf '\033[1;36m[deploy]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[deploy][warn]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[deploy][err]\033[0m %s\n' "$*" >&2; exit 1; }
hr() { printf '\033[2m%s\033[0m\n' "----------------------------------------------------------------"; }
# Return the host-published TCP ports this deployment type will bind, so we can
# check them for collisions BEFORE bring-up. Ports come from docker-compose.yml
# ("HOST:CONTAINER" left side). Kept explicit here (not parsed) so the check is
# predictable; keep in sync with compose if published ports change.
# full-stack / central-store: full UI+store surface. remote: just BMP+kafka.
ports_for_type() {
case "$1" in
remote)
# BMP in (router-facing) + kafka out
echo "$ROUTER_PORT 9092" ;;
central-store)
# no local collector -> no BMP listen; store+UI+bus
echo "9092 5432 3000 4300" ;;
full-stack|*)
# BMP in + kafka + postgres + grafana + whois
echo "$ROUTER_PORT 9092 5432 3000 4300" ;;
esac
}
# What is a port MAPPED to internally, for the message (router-facing -> 5000).
port_note() {
case "$1" in
"$ROUTER_PORT") echo "BMP (maps to collector :$COLLECTOR_PORT)" ;;
9092) echo "Kafka" ;;
5432) echo "Postgres" ;;
3000) echo "Grafana" ;;
4300) echo "whois" ;;
*) echo "stack service" ;;
esac
}
# Check each port for an existing listener on the host. Uses ss, falls back to
# nc/netstat. Warns per-collision (with what's likely holding it) and returns
# 1 if any collision found. Never hard-fails on its own - caller decides.
preflight_ports() {
local type="$1" p conflict=0 tool=""
if command -v ss >/dev/null 2>&1; then tool="ss"
elif command -v netstat >/dev/null 2>&1; then tool="netstat"
elif [ -r /proc/net/tcp ] || [ -r /proc/net/tcp6 ]; then tool="proc"
else
warn "No ss/netstat//proc net available - skipping port-collision check."
return 0
fi
# Build the set of listening ports once (as decimal), for the proc path.
local listening=""
if [ "$tool" = "proc" ]; then
# /proc/net/tcp{,6}: col 2 = local_addr:PORThex, col 4 = state (0A=LISTEN).
# Convert the hex port to decimal WITHOUT gawk's strtonum (mawk lacks it).
listening="$(cat /proc/net/tcp /proc/net/tcp6 2>/dev/null \
| awk '$4=="0A"{n=split($2,a,":"); print a[n]}' \
| while read -r hx; do printf '%d\n' "0x$hx"; done | sort -u)"
fi
port_busy() {
local port="$1"
case "$tool" in
ss) ss -ltn 2>/dev/null | awk '{print $4}' | grep -qE "[:.]${port}\$" ;;
netstat) netstat -ltn 2>/dev/null | awk '{print $4}' | grep -qE "[:.]${port}\$" ;;
proc) grep -qx "$port" <<<"$listening" ;;
esac
}
for p in $(ports_for_type "$type"); do
if port_busy "$p"; then
conflict=1
local who=""
if command -v ss >/dev/null 2>&1; then
who="$(ss -ltnp 2>/dev/null | grep -E "[:.]${p}\b" | grep -oE 'users:\(\("[^"]+"' | head -1 | sed 's/users:((//; s/"//g')"
fi
warn "Port ${p} ($(port_note "$p")) is ALREADY IN USE${who:+ by: $who}"
fi
done
return $conflict
}
# Prompt with an editable default. Honors --yes (accepts default silently).
ask() {
local __var="$1" __prompt="$2" __default="${3:-}" __reply=""
if [ "$ARG_YES" -eq 1 ]; then
printf -v "$__var" '%s' "$__default"; return
fi
if [ -n "$__default" ]; then
read -r -p "$__prompt [$__default]: " __reply
__reply="${__reply:-$__default}"
else
read -r -p "$__prompt: " __reply
fi
printf -v "$__var" '%s' "$__reply"
}
confirm() {
local __prompt="$1" __reply=""
[ "$ARG_YES" -eq 1 ] && return 0
read -r -p "$__prompt [y/N]: " __reply
[ "$__reply" = "y" ] || [ "$__reply" = "Y" ]
}
cd "$REPO_DIR" || die "REPO_DIR '$REPO_DIR' not found"
[ -f docker-compose.yml ] || die "no docker-compose.yml in $REPO_DIR - set REPO_DIR"
command -v docker >/dev/null || die "docker not found"
docker compose version >/dev/null 2>&1 || die "docker compose v2 not available"
# --- Docker Desktop vs native-daemon guard ----------------------------------
# On WSL, Docker Desktop runs the daemon in its OWN separate distro
# (docker-desktop VM). Host bind mounts (this stack binds ${OBMP_DATA_ROOT})
# then resolve against THAT VM's filesystem, not this distro's - so a dir that
# exists in `ls` fails to mount with a misleading "no such file or directory".
# Detect it and warn loudly, because the failure is silent and wastes a deploy.
detect_docker_desktop() {
# Fast signals first (no daemon round-trip): the Desktop proxy process and
# the desktop-linux context.
if ps -eo args 2>/dev/null | grep -q '[d]ocker-desktop.*proxy'; then return 0; fi
if docker context inspect 2>/dev/null | grep -qi 'desktop-linux\|dockerDesktop'; then return 0; fi
# Authoritative: the daemon reports itself.
docker info 2>/dev/null | grep -qiE 'Operating System:.*Docker Desktop|Name:.*docker-desktop'
}
if detect_docker_desktop; then
warn "Docker Desktop detected as the active daemon."
warn "This stack uses HOST BIND MOUNTS under \$OBMP_DATA_ROOT. Under Docker"
warn "Desktop's WSL integration the daemon runs in a SEPARATE distro, so"
warn "those binds resolve against the wrong filesystem and fail to mount"
warn "with a misleading 'no such file or directory' - even though the dir"
warn "exists in your shell. Use NATIVE docker-in-WSL instead:"
warn " 1) Docker Desktop -> Settings -> Resources -> WSL Integration ->"
warn " turn OFF integration for this distro"
warn " 2) install docker-ce in this distro and: sudo systemctl enable --now docker"
warn " 3) confirm: docker info | grep -i 'operating system' (should NOT say Docker Desktop)"
if [ "$ARG_YES" -eq 1 ]; then
die "refusing to deploy onto Docker Desktop with host bind mounts under --yes"
fi
if ! confirm "Continue anyway (bind mounts will likely fail)?"; then
die "aborted - switch to native docker-in-WSL and re-run"
fi
fi
# --- .env helpers -----------------------------------------------------------
if [ ! -f .env ]; then
[ -f .env.example ] || die "no .env or .env.example present"
cp .env.example .env
log "Created .env from .env.example"
fi
get_env() { grep -E "^$1=" .env | head -1 | cut -d= -f2- || true; }
set_env() {
local key="$1" val="$2"
if grep -qE "^${key}=" .env; then
sed -i "s|^${key}=.*|${key}=${val}|" .env
else
printf '%s=%s\n' "$key" "$val" >> .env
fi
}
# --- 1. host type: auto-detect, then override -------------------------------
detected="prod"
if grep -qiE 'microsoft|wsl' /proc/version 2>/dev/null; then detected="wsl"; fi
if [ -n "$ARG_HOSTTYPE" ]; then
HOSTTYPE="$ARG_HOSTTYPE"
log "Host type: $HOSTTYPE (from flag; auto-detect said '$detected')"
else
log "Auto-detected host type: $detected"
ask HOSTTYPE "Host type (wsl/prod)" "$detected"
fi
case "$HOSTTYPE" in wsl|prod) ;; *) die "host type must be 'wsl' or 'prod' (got '$HOSTTYPE')";; esac
# --- 2. deployment TYPE (described menu, resolved before auth/endpoints) -----
# Three types from docs/DEPLOYMENT-TYPES.md. Old scope names still accepted as
# aliases via --scope so existing invocations don't break.
# full-stack = collocated ingest + store (the old 'feeders'/'full')
# remote = collector/forwarder only (the old 'standalone')
# central-store = store only, no local collector (new)
print_type_menu() {
cat <<'EOF'
Deployment types (what runs here / where data lands):
1) full-stack Everything on one host: collector + Kafka + Postgres +
Grafana + feeders. BMP terminates and data is stored here.
Use when one node monitors the whole fabric.
RESOURCES: prod-realistic 16 vCPU / 48-64 GB / NVMe >=250 GB;
lab-only 4 vCPU / 16 GB host / SSD.
2) remote Collector + local Kafka only, FORWARDS to a central store's
Kafka. No local Postgres/Grafana. Run near the routers to
spread ingest / distribute BMP session load.
RESOURCES: light - 2-4 vCPU / 4-8 GB / 20-40 GB fast disk
(Kafka spool only). Needs --central-kafka HOST:PORT.
3) central-store Store half: Kafka + Postgres + Grafana + feeders, NO local
collector. Ingest arrives from remote nodes over Kafka.
Run one of these behind N remote collectors.
RESOURCES: same as full-stack store tier - 16 vCPU /
48-64 GB / NVMe >=250 GB (carries all remotes' data).
Note: resource figures are engineering estimates to calibrate against your
own watermarking, not measured specs. Authelia auth is a SEPARATE toggle.
EOF
}
# normalize any alias to a canonical type
normalize_type() {
case "$1" in
full-stack|full_stack|fullstack|feeders|full|core) echo "full-stack" ;;
remote|remote-collector|standalone|collector|forwarder) echo "remote" ;;
central-store|central|store) echo "central-store" ;;
*) echo "" ;;
esac
}
if [ -n "$ARG_SCOPE" ]; then
DTYPE="$(normalize_type "$ARG_SCOPE")"
[ -n "$DTYPE" ] || die "unknown --scope '$ARG_SCOPE' (use full-stack|remote|central-store)"
else
default_type="full-stack"
if [ "$ARG_YES" -eq 0 ]; then print_type_menu; fi
_t="$default_type"
ask _t "Deployment type (full-stack/remote/central-store)" "$default_type"
DTYPE="$(normalize_type "$_t")"
[ -n "$DTYPE" ] || die "unknown type '$_t' (use full-stack|remote|central-store)"
fi
log "Deployment type: $DTYPE"
# --- 3. HOST_IP: smart default, editable, validated -------------------------
# HOST_IP is the address the INTERNAL stack advertises/binds (Kafka listener,
# gobgp->collector). On a remote collector it is also this node's own address.
primary_ip="$(ip -o -4 addr show scope global 2>/dev/null | awk '{print $4}' | cut -d/ -f1 | head -1 || true)"
wsl_ip="$(hostname -I 2>/dev/null | awk '{print $1}' || true)"
current_ip="$(get_env HOST_IP)"
if [ "$HOSTTYPE" = "wsl" ]; then
default_ip="${ARG_HOST_IP:-${wsl_ip:-$current_ip}}"
else
default_ip="${ARG_HOST_IP:-${primary_ip:-$current_ip}}"
fi
if [ -n "$ARG_HOST_IP" ]; then
HOST_IP="$ARG_HOST_IP"
else
if [ -n "$current_ip" ] && [ "$current_ip" != "$default_ip" ]; then
warn "Existing HOST_IP in .env is '$current_ip' - may be stale (carried from another host)."
fi
ask HOST_IP "HOST_IP for the internal stack (Kafka listener / collector bind)" "$default_ip"
fi
[ -n "$HOST_IP" ] || die "HOST_IP cannot be empty"
local_addrs="$(ip -o addr show 2>/dev/null | awk '{print $4}' | cut -d/ -f1 || true)"
if [ -z "$local_addrs" ]; then
warn "Could not enumerate local addresses ('ip' missing?) - skipping HOST_IP validation."
elif ! grep -qx "$HOST_IP" <<<"$local_addrs"; then
if [ "$HOSTTYPE" = "wsl" ]; then
warn "HOST_IP $HOST_IP is not a local addr (expected on WSL; NAT/portproxy applies)."
else
warn "HOST_IP $HOST_IP is NOT a local address on this host. Local addresses:"
ip -o addr show | awk '{print " "$2" "$4}' | grep -v '127.0.0.1' >&2 || true
warn "This is the classic ported-.env bug. Only proceed if a VIP/NAT forwards to this host."
fi
fi
# --- 4. router-facing IP + port (what routers actually target for BMP) ------
# This is deliberately SEPARATE from HOST_IP:
# - On WSL, routers cannot reach the WSL NAT addr; they hit the Windows LAN
# IP, forwarded by netsh portproxy into the WSL collector.
# - On prod, routers usually reach HOST_IP directly, but a VIP/NAT/jump can
# make the router-facing address differ from the bind address.
# It feeds the printed portproxy commands and post-deploy guidance. The
# collector's own listen port stays 5000 inside the container regardless.
COLLECTOR_PORT=5000 # in-container BMP listen port (fixed by the image)
# Best-effort: from inside WSL, ask the Windows host for its LAN IP via
# powershell.exe interop. Returns empty if not WSL, powershell.exe is absent,
# or nothing plausible is found. Filters out the WSL vEthernet, loopback,
# link-local/APIPA (169.254), and picks the interface that owns the default
# route (the real LAN NIC) when possible.
detect_windows_lan_ip() {
command -v powershell.exe >/dev/null 2>&1 || return 0
# Preferred IPv4 = the source address of the default route on Windows.
local ps_out
ps_out="$(powershell.exe -NoProfile -Command \
"(Get-NetIPConfiguration | Where-Object { \$_.IPv4DefaultGateway -ne \$null -and \$_.NetAdapter.Status -eq 'Up' } | Select-Object -First 1 -ExpandProperty IPv4Address).IPAddress" \
2>/dev/null | tr -d '\r' | tr -d '[:space:]')"
# Fallback: first non-WSL, non-APIPA IPv4 on any up adapter.
if [ -z "$ps_out" ]; then
ps_out="$(powershell.exe -NoProfile -Command \
"Get-NetIPAddress -AddressFamily IPv4 | Where-Object { \$_.InterfaceAlias -notlike '*WSL*' -and \$_.IPAddress -notlike '169.254.*' -and \$_.IPAddress -ne '127.0.0.1' } | Select-Object -First 1 -ExpandProperty IPAddress" \
2>/dev/null | tr -d '\r' | tr -d '[:space:]')"
fi
# sanity: looks like a dotted quad
case "$ps_out" in
[0-9]*.[0-9]*.[0-9]*.[0-9]*) printf '%s' "$ps_out" ;;
*) return 0 ;;
esac
}
if [ "$HOSTTYPE" = "wsl" ]; then
win_lan_ip=""
if [ -z "$ARG_ROUTER_IP" ]; then
win_lan_ip="$(detect_windows_lan_ip || true)"
if [ -n "$win_lan_ip" ]; then
log "Detected Windows host LAN IP: $win_lan_ip (routers reach the stack here)"
else
warn "Could not auto-detect the Windows LAN IP - falling back to a placeholder."
warn "Find it on Windows with: Get-NetIPConfiguration (use the 'Up' adapter's IPv4Address)"
fi
fi
default_router_ip="${ARG_ROUTER_IP:-${win_lan_ip:-<WINDOWS_LAN_IP>}}"
else
default_router_ip="${ARG_ROUTER_IP:-$HOST_IP}"
fi
# The router-facing/external BMP port defaults to 1790 (the IANA-registered BMP
# port) rather than the collector's internal 5000 - avoids the crowded 5000 and
# is self-documenting. The collector still listens on 5000 INSIDE the container;
# compose publishes ${ROUTER_FACING_PORT:-1790}:5000 everywhere, so on WSL the
# portproxy forwards listenport=ROUTER_PORT to the SAME port on the WSL side.
BMP_STD_PORT=1790
default_router_port="${ARG_ROUTER_PORT:-$BMP_STD_PORT}"
if [ -n "$ARG_ROUTER_IP" ]; then ROUTER_IP="$ARG_ROUTER_IP"; else
ask ROUTER_IP "Router-facing IP (what 'bmp server' on the routers targets)" "$default_router_ip"
fi
if [ -n "$ARG_ROUTER_PORT" ]; then ROUTER_PORT="$ARG_ROUTER_PORT"; else
ask ROUTER_PORT "Router-facing BMP port" "$default_router_port"
fi
[ -n "$ROUTER_IP" ] || die "router-facing IP cannot be empty"
[ -n "$ROUTER_PORT" ] || die "router-facing port cannot be empty"
# --- 5. remote: central Kafka endpoint --------------------------------------
CENTRAL_KAFKA=""
if [ "$DTYPE" = "remote" ]; then
default_ck="${ARG_CENTRAL_KAFKA:-$(get_env KAFKA_FQDN)}"
if [ -n "$ARG_CENTRAL_KAFKA" ]; then CENTRAL_KAFKA="$ARG_CENTRAL_KAFKA"; else
ask CENTRAL_KAFKA "Central store Kafka endpoint (HOST:PORT)" "$default_ck"
fi
[ -n "$CENTRAL_KAFKA" ] || die "remote type requires a central Kafka endpoint (--central-kafka HOST:PORT)"
case "$CENTRAL_KAFKA" in *:*) ;; *) die "central Kafka must be HOST:PORT (got '$CENTRAL_KAFKA')";; esac
fi
# --- 6. Authelia toggle (independent axis; N/A for remote - no local UI) -----
AUTH_MODE="local"
OBMP_DOMAIN="$(get_env OBMP_DOMAIN)"
OBMP_COOKIE_DOMAIN="$(get_env OBMP_COOKIE_DOMAIN)"
if [ "$DTYPE" = "remote" ]; then
AUTH_MODE="local" # remote collector runs no Grafana/portal
else
current_auth="$(get_env OBMP_AUTH_MODE)"
if [ -n "$ARG_AUTH" ]; then
AUTH_MODE="$ARG_AUTH"
else
default_auth="local"
[ "$HOSTTYPE" = "prod" ] && default_auth="${current_auth:-local}"
if [ "$ARG_YES" -eq 0 ]; then
printf '\n Authelia = SSO/reverse-proxy in front of Grafana (needs a real domain).\n local = Grafana served directly on :3000 (fine for lab/test).\n'
fi
ask AUTH_MODE "Front Grafana with Authelia? (local/authelia)" "$default_auth"
fi
case "$AUTH_MODE" in local|authelia) ;; *) die "auth mode must be 'local' or 'authelia' (got '$AUTH_MODE')";; esac
if [ "$AUTH_MODE" = "authelia" ]; then
if [ -z "$OBMP_DOMAIN" ] || [[ "$OBMP_DOMAIN" == changeme* ]] || [[ "$OBMP_DOMAIN" == nowhere.example.com ]]; then
ask OBMP_DOMAIN "Public domain fronting Grafana/Authelia" "${OBMP_DOMAIN}"
[ -n "$OBMP_DOMAIN" ] && [[ "$OBMP_DOMAIN" != nowhere.example.com ]] \
|| die "authelia mode needs a real OBMP_DOMAIN"
fi
ask OBMP_COOKIE_DOMAIN "Authelia cookie domain (parent of above)" "${OBMP_COOKIE_DOMAIN:-$OBMP_DOMAIN}"
fi
fi
OBMP_DATA_ROOT="$(get_env OBMP_DATA_ROOT)"; OBMP_DATA_ROOT="${OBMP_DATA_ROOT:-/var/openbmp}"
# --- 7. PLAN ----------------------------------------------------------------
type_desc() {
case "$1" in
full-stack) echo "collector + Kafka + Postgres + Grafana + feeders (ingest + store on one host)" ;;
remote) echo "collector + local Kafka only -> forwards to central Kafka (no local store)" ;;
central-store) echo "Kafka + Postgres + Grafana + feeders, NO local collector (ingest via remotes)" ;;
esac
}
type_resources() {
case "$1" in
full-stack) echo "prod ~16 vCPU / 48-64 GB / NVMe >=250 GB (lab 4 vCPU / 16 GB / SSD)" ;;
remote) echo "~2-4 vCPU / 4-8 GB / 20-40 GB fast disk (Kafka spool only)" ;;
central-store) echo "~16 vCPU / 48-64 GB / NVMe >=250 GB (carries all remotes' data)" ;;
esac
}
echo; hr; log "Deployment plan"; hr
# Record what is being deployed - a greenfield deploy is only reproducible if
# the checkout is pinned (branch + commit go in the plan and the terminal log).
git_desc="$(git -C "$REPO_DIR" rev-parse --abbrev-ref HEAD 2>/dev/null || true)"
[ -n "$git_desc" ] && git_desc="$git_desc @ $(git -C "$REPO_DIR" rev-parse --short HEAD 2>/dev/null)$(git -C "$REPO_DIR" diff --quiet 2>/dev/null || echo ' (dirty)')"
cat <<EOF
Repo dir : $REPO_DIR
Checkout : ${git_desc:-not a git checkout}
Host type : $HOSTTYPE (auto-detected: $detected)
Deployment type : $DTYPE
$(type_desc "$DTYPE")
Resources (est) : $(type_resources "$DTYPE")
Data root : $OBMP_DATA_ROOT
HOST_IP : $HOST_IP (internal stack bind/advertise)
Router-facing : ${ROUTER_IP}:${ROUTER_PORT} (what 'bmp server' targets)
EOF
[ "$DTYPE" = "remote" ] && echo " Central Kafka : $CENTRAL_KAFKA"
[ "$DTYPE" != "remote" ] && printf ' Grafana auth : %s%s\n' "$AUTH_MODE" "$([ "$AUTH_MODE" = authelia ] && printf ' (domain: %s)' "$OBMP_DOMAIN")"
echo " Reset data : $([ "$ARG_RESET" -eq 1 ] && echo 'YES - destroys data tree' || echo 'no')"
echo
echo " Will run setup.sh (provision data tree, render configs, pull/build)"
echo " Then: clear stale volume handles -> staged bring-up -> verify"
hr
# --- port-collision preflight (before touching anything) --------------------
if preflight_ports "$DTYPE"; then
: # no collisions
else
echo
warn "One or more host ports the stack needs are already in use (see above)."
warn "If you continue, those containers will FAIL to bind and come up unhealthy."
warn "Options: stop the conflicting service, or re-run choosing different"
warn "published ports (e.g. --router-port for BMP; Grafana/Postgres/Kafka"
warn "ports are set in docker-compose.yml)."
if [ "$ARG_YES" -eq 1 ]; then
die "refusing to proceed under --yes with a port collision; resolve it or change ports"
fi
if ! confirm "Proceed ANYWAY despite the port collision?"; then
die "aborted due to port collision - nothing was modified"
fi
fi
if ! confirm "Proceed with this plan?"; then
die "aborted before making changes - nothing was modified"
fi
# --- 8. write .env ----------------------------------------------------------
set_env HOST_IP "$HOST_IP"
set_env OBMP_AUTH_MODE "$AUTH_MODE"
if [ "$AUTH_MODE" = "authelia" ]; then
set_env OBMP_DOMAIN "$OBMP_DOMAIN"
set_env OBMP_COOKIE_DOMAIN "$OBMP_COOKIE_DOMAIN"
fi
# Persist router-facing target for reference/other tooling.
set_env ROUTER_FACING_IP "$ROUTER_IP"
set_env ROUTER_FACING_PORT "$ROUTER_PORT"
# Remote: point the collector at the central bus.
if [ "$DTYPE" = "remote" ]; then
set_env KAFKA_FQDN "$CENTRAL_KAFKA"
fi
log ".env reconciled for this host"
# --- 9. optional reset ------------------------------------------------------
if [ "$ARG_RESET" -eq 1 ]; then
if [ "$HOSTTYPE" = "prod" ]; then
warn "PROD reset will DELETE ALL DATA under $OBMP_DATA_ROOT (incl. Postgres BMP history)."
[ "$ARG_YES" -eq 1 ] && die "refusing unattended --reset on prod; run interactively to confirm"
read -r -p "Type the data root path to confirm deletion: " __c
[ "$__c" = "$OBMP_DATA_ROOT" ] || die "confirmation did not match - aborted"
else
warn "Resetting data tree under $OBMP_DATA_ROOT (WSL test host)"
fi
docker compose --profile test --profile auth down -v 2>/dev/null || true
docker volume rm obmp_data-volume obmp_ts-volume 2>/dev/null || true
sudo rm -rf "${OBMP_DATA_ROOT:?}/"* 2>/dev/null || true
log "Reset complete"
fi
# --- 10. run setup.sh -------------------------------------------------------
# Clear any stale state BEFORE setup.sh, so a "failed to populate" volume
# handle from a prior aborted run can't survive into this deployment. Bind
# volumes (type:none,o:bind) cache a mount-failure state that replays even
# after the source dir exists - removing the handle forces a fresh bind.
docker compose down 2>/dev/null || true
docker volume rm obmp_data-volume obmp_ts-volume 2>/dev/null || true
[ -x ./setup.sh ] || chmod +x ./setup.sh 2>/dev/null || true
[ -f ./setup.sh ] || die "setup.sh not found in $REPO_DIR - cannot provision"
log "Running setup.sh ..."
./setup.sh || die "setup.sh failed - fix the reported issue and re-run"
# Verify setup.sh actually produced the bind-mount source dirs compose needs.
# If these are missing, every psql 'up' fails with 'no such file or directory'.
_dr="$(get_env OBMP_DATA_ROOT)"; _dr="${_dr:-/var/openbmp}"
for _d in postgres/data postgres/ts; do
if [ ! -d "$_dr/$_d" ]; then
warn "Expected bind dir $_dr/$_d is missing after setup.sh - creating it."
sudo mkdir -p "$_dr/$_d" || die "could not create $_dr/$_d"
fi
done
# Clear the volume handles again right before bring-up (setup.sh's pull/build
# or any concurrent action could have recreated them).
docker compose down 2>/dev/null || true
docker volume rm obmp_data-volume obmp_ts-volume 2>/dev/null || true
# --- 11. staged bring-up ----------------------------------------------------
wait_kafka() {
log " infra (zookeeper, kafka) ..."
docker compose up -d zookeeper kafka
for _ in $(seq 1 30); do
[ "$(docker inspect obmp-kafka --format '{{.State.Status}}' 2>/dev/null || echo x)" = running ] && break
sleep 2
done
[ "$(docker inspect obmp-kafka --format '{{.State.Status}}' 2>/dev/null)" = running ] \
|| die "kafka did not stabilise - check: docker logs obmp-kafka"
log " kafka up"
}
wait_postgres() {
log " waiting for postgres ..."
for _ in $(seq 1 60); do
docker exec obmp-psql pg_isready -U openbmp -d openbmp >/dev/null 2>&1 && break
sleep 3
done
}
if [ "$DTYPE" = "remote" ]; then
# Remote collector: ingest path only, forwards to central Kafka.
log "Stage 1/2 - infra"
wait_kafka
log "Stage 2/2 - collector (forwarding to $CENTRAL_KAFKA) ..."
docker compose up -d collector
log " collector up"
elif [ "$DTYPE" = "central-store" ]; then
# Store only: Postgres/psql-app/Grafana/feeders, NO local collector.
log "Stage 1/3 - infra"
wait_kafka
log "Stage 2/3 - store (psql, psql-app, grafana, whois) ..."
docker compose up -d psql
wait_postgres
docker compose up -d psql-app grafana whois
log " store up"
log "Stage 3/3 - feeders/consumers (--profile test) ..."
docker compose --profile test up -d
# ensure no local collector is running on a central-store node
docker compose stop collector 2>/dev/null || true
log " feeders up (no local collector)"
else
# full-stack: collector + store + feeders on one host.
profiles=(--profile test)
[ "$AUTH_MODE" = "authelia" ] && profiles+=(--profile auth)
log "Stage 1/3 - infra"
wait_kafka
log "Stage 2/3 - core (psql, collector, psql-app, grafana, whois) ..."
docker compose up -d psql
wait_postgres
docker compose up -d collector psql-app grafana whois
log " core up"
log "Stage 3/3 - feeders${AUTH_MODE:+ + auth if authelia} (${profiles[*]}) ..."
docker compose "${profiles[@]}" up -d
log " feeders up"
fi
# --- 12. verify -------------------------------------------------------------
echo; log "Current state:"
docker compose ps
echo; log "Collector -> Kafka (want NO 'Failed to resolve' / 'brokers down'):"
docker logs obmp-collector --tail 8 2>&1 | sed 's/^/ /' || true
# --- 13. post-deploy notes --------------------------------------------------
echo
printf '\033[1;32m[deploy] done\033[0m\n'
if [ "$DTYPE" = "remote" ]; then
cat <<EOF
Remote collector active. BMP ingest -> $CENTRAL_KAFKA (central store).
This node runs NO local Postgres/Grafana; view data on the central node.
Point routers' 'bmp server' at ${ROUTER_IP}:${ROUTER_PORT}.
EOF
elif [ "$DTYPE" = "central-store" ]; then
if [ "$AUTH_MODE" = "authelia" ]; then
echo " Grafana: https://${OBMP_DOMAIN}/grafana/"
else
echo " Grafana: http://${HOST_IP}:3000/grafana/"
fi
echo " Central store active. Ingest arrives from remote collectors over Kafka."
echo " No local collector here - routers point at the REMOTE nodes, not this host."
else
if [ "$AUTH_MODE" = "authelia" ]; then
echo " Grafana: https://${OBMP_DOMAIN}/grafana/"
else
echo " Grafana: http://${HOST_IP}:3000/grafana/"
fi
echo " BMP collector: routers target ${ROUTER_IP}:${ROUTER_PORT}"
fi
if [ "$HOSTTYPE" = "wsl" ] && [ "$DTYPE" != "central-store" ]; then
cat <<EOF
WSL note: HOST_IP ${HOST_IP} is the WSL VM address - reachable from Windows
and the compose network, but NOT from physical routers, and it changes on
'wsl --shutdown'. Routers must target the router-facing address
(${ROUTER_IP}:${ROUTER_PORT}) via a portproxy on the Windows host. In an
ELEVATED PowerShell, run the repo's helper (auto-detects the current WSL IP,
idempotent):
powershell -ExecutionPolicy Bypass -File scripts/wsl-portproxy.ps1 -RouterPort ${ROUTER_PORT}
or add the rules by hand:
netsh interface portproxy add v4tov4 listenport=${ROUTER_PORT} listenaddress=0.0.0.0 connectport=${ROUTER_PORT} connectaddress=${HOST_IP}
netsh interface portproxy add v4tov4 listenport=9092 listenaddress=0.0.0.0 connectport=9092 connectaddress=${HOST_IP}
New-NetFirewallRule -DisplayName "OpenBMP ${ROUTER_PORT}" -Direction Inbound -LocalPort ${ROUTER_PORT} -Protocol TCP -Action Allow
Re-run after any WSL restart (the WSL IP changes). See docs/router-bmp-config.md.
EOF
elif [ "$DTYPE" != "central-store" ]; then
cat <<EOF
Native host. Ensure the firewall permits inbound TCP ${ROUTER_PORT} from the
router mgmt network. Router config: see docs/router-bmp-config.md (use
${ROUTER_IP}:${ROUTER_PORT} as the bmp server host/port).
EOF
fi