================================================================================ RODMENA PLATFORM TOPOLOGY ================================================================================ Last verified : 2026-10-05 Scope : the first-party platform services and the data tier behind them Audience : engineers integrating with, operating, or auditing the platform This document describes architecture. It omits operational coordinates (host addresses, database endpoints, role names, storage bucket names, internal ports, filesystem paths). See NOTES ON DISCLOSURE at the end. ================================================================================ 1. AT A GLANCE ================================================================================ Managed database fleet: 28 PostgreSQL databases on 4 database hosts (3 primaries, 1 standby) in 3 countries Of the 28, 22 are production databases, each belonging to exactly one service. The remaining 6 are development, test and pre-production copies, plus one copy being retired after a move between regions. They are not enumerated here, and they are not a lower tier of care: every database on these hosts is covered by the certificate, isolation, replication and backup regime in sections 5 to 7, because a development copy holds real shapes and sometimes real data. A small number of services keep their database on their application host instead of in the managed fleet. See "Databases outside the fleet" below. Every service is a separate application with its own database. No two services share a schema, and no service reads another's tables. All cross-service communication is over HTTPS APIs. Applications and data are kept in separate places, because only the applications can be redeployed. APPLICATION TIER DATA TIER CONSENSUS TIER United Kingdom United Kingdom / United Kingdom / France / Germany France / Germany (4 hosts, 3 regions) (3 members, 3 regions) The application tier is entirely in the United Kingdom, on three kinds of host: virtual machines provided by a UK hosting company (iDNet), UK virtual servers from OVH, and UK virtual servers from Cloud Nord. A few services run on company-owned machines at RODMENA's office. Application hosts hold no production data of the fleet's services; their data is in the data tier. Databases outside the fleet. 4 production databases live on an application host instead of in the managed fleet: those of the status-page service, the tunnel service, webmail and one further workload. They are not covered by sections 6 and 7. They have no standby replica, and they are not in the continuous backup system. They are dumped nightly, and the dumps are included in a nightly encrypted snapshot to off-site object storage. Their recovery point is therefore about 24 hours, not 5 minutes, and recovery is a restore from the most recent nightly dump. ================================================================================ 2. SERVICES ================================================================================ Each entry: what it does, where its data lives, and what it calls at runtime. "Region" refers to the data centre holding that service's database. -------------------------------------------------------------------------------- AUTH https://auth.rodmena.co.uk -------------------------------------------------------------------------------- Role Role-based access control and authorisation for the platform Database Region FR Depends on (nothing) Depended on by futex, mail_api, agentbus, red9, runflow Auth is the root of the trust tree. It calls no other first-party service, because an authorisation service that depends on its own consumers cannot be recovered independently. -------------------------------------------------------------------------------- TOKENGATE https://tokengate.rodmena.co.uk -------------------------------------------------------------------------------- Role Quota management, rate limiting and usage metering as a service Database Region FR Depends on (nothing) Depended on by mail_api, identity, agentbus TokenGate is stateless-recoverable. Its hot-path counters are held in a cache and are rebuilt from the database of record if that cache is lost. -------------------------------------------------------------------------------- IDENTITY https://identity.rodmena.co.uk -------------------------------------------------------------------------------- Role OpenID Connect identity provider for end users; issues and rotates signing keys, manages sessions and consent Database Region FR Depends on mail_api, tokengate Depended on by red9 Identity and auth are separate systems with separate databases. Identity authenticates people; auth authorises services. -------------------------------------------------------------------------------- FUTEX https://futex.rodmena.co.uk -------------------------------------------------------------------------------- Role Human-in-the-loop approvals: pauses automated work pending a human decision, with policies and audit Database Region UK Depends on auth, mail_api, runflow Depended on by agentbus, red9 -------------------------------------------------------------------------------- RUNFLOW https://runflow.rodmena.co.uk -------------------------------------------------------------------------------- Role Workflow and pipeline engine; scheduled and event-driven execution in isolated containers Database Region UK Depends on auth, tokengate, mail_api Depended on by futex, mail_api By transaction volume this is the busiest database on the platform. ISOLATED HOST. RunFlow is the only service that executes untrusted, caller-supplied container workloads. It therefore runs on a dedicated host in the application tier and shares that host with nothing else. A compromise of a caller's container reaches no other service's process, configuration or credentials. The host holds a pull-only credential for the container registry, so an escape cannot publish an image that other services would later run. Container network egress is filtered. Callers' containers may reach the public internet, which is an advertised capability. They are refused every private range and every first-party host, including the database tier. First-party hosts are on public addresses, so access to them has to be denied explicitly. Both names resolve to that host: runflow.rodmena.co.uk is primary and runflow.rodmena.app is retained until callers migrate. A SECOND, INDEPENDENT INSTANCE runs as a UK vantage node on its own small host, with its own database, its own scheduler lock and its own credentials. It is deliberately not a failover target and not a capacity extension. It is sized for scheduled outbound checks from a different network, so that an availability measurement is not always taken from the same place. The two instances share nothing -- not a database, not a lock, not a host, not a network -- so the loss of either leaves the other unaffected. -------------------------------------------------------------------------------- MAIL API https://mailserver.rodmena.co.uk -------------------------------------------------------------------------------- Role Transactional and campaign email: sending, delivery tracking, inbound routing, suppression and complaint handling Database Region UK Depends on auth, runflow, tokengate Depended on by identity, futex, agentbus, red9, runflow, ledger Mail API is the most depended-upon service after auth. Attachments are stored on the application host, outside the database. -------------------------------------------------------------------------------- AGENTBUS https://agentbus.rodmena.co.uk -------------------------------------------------------------------------------- Role Inter-platform message bus: addressed, threaded messaging between autonomous agents, with delivery guarantees and audit Database Region UK Depends on auth, futex, mail_api, tokengate Depended on by (operational tooling) AgentBus runs as eight coordinated processes: two API instances behind a load balancer, plus egress, ingest, webhook, janitor and synthetic-check workers. -------------------------------------------------------------------------------- RED9 (not yet in service) -------------------------------------------------------------------------------- Status Not running. Coming soon; the description below is the intended design. Role Multi-tenant task and conversation platform Database Region UK Depends on auth, identity, futex, mail_api, and an external model provider Depended on by (none) Red9 is the only service that enforces tenant isolation inside the database, using row-level security keyed on a per-request tenant identifier. Its application role cannot bypass those policies. -------------------------------------------------------------------------------- LEDGER https://ledger.rodmena.com -------------------------------------------------------------------------------- Role Double-entry accounting ledger: assets, postings, balances Database Region UK Depends on mail_api Depended on by (none) Enforces per-tenant isolation with row-level security and separates the application, maintenance and reporting roles so that ordinary request handling cannot perform maintenance operations. -------------------------------------------------------------------------------- KNOWLEDGE BASE https://kb.rodmena.co.uk -------------------------------------------------------------------------------- Role Documentation platform: workspaces, spaces, a tree of versioned Markdown documents, attachments, search and share links Database Region UK Depends on auth, identity, tokengate, mail_api, futex, runflow Depended on by every service, as the system of record for its operational documentation IN SERVICE since September 2026. An earlier revision of this document listed it as pre-production; it now carries the estate's operational documentation and is the source of record for it. Document versions are immutable, and a published pointer selects which version a reader sees, so a page cannot be rewritten under someone reading it. Attachments are encrypted at rest and served through the API rather than from a storage URL. Machine callers hold service-account keys scoped to one workspace and one set of spaces. Governance actions -- audit, retention, membership, purge, legal holds -- are refused to every machine credential whatever its role, and require a signed-in person. A credential outside its scope is answered with "not found" rather than "forbidden", so a narrowed key cannot map what it cannot reach. -------------------------------------------------------------------------------- PROVENANCE https://provenance.rodmena.co.uk -------------------------------------------------------------------------------- Role Software supply-chain provenance: SBOM ingest and component inventory, dependency scanning, signed report export Database Region FR Depends on auth, tokengate Depended on by engineering tooling Keeps a separate development database alongside production, on the same host under the same controls, so a schema change can be exercised against real shapes without touching the production copy. -------------------------------------------------------------------------------- TRACE https://trace.rodmena.co.uk -------------------------------------------------------------------------------- Role Execution tracing and planning records for automated work Database Region UK Depends on auth Depended on by engineering tooling -------------------------------------------------------------------------------- VELLUM https://vellum.rodmena.co.uk -------------------------------------------------------------------------------- Role Subject Access Request management and redaction: intake, rendition, redaction, and production of disclosure bundles Database Region FR Depends on auth, tokengate Depended on by (none) This service handles the most sensitive category of data on the platform. The object-storage path is DESIGNED so that files are encrypted CLIENT-SIDE before they leave the application host -- a per-object symmetric key wrapped to an asymmetric recipient key -- so the object store and its operator hold ciphertext only. Case files are held in a UK object-storage region, separate from the platform's other storage, under a credential that can reach no other container, cannot delete the container itself, and cannot alter its retention rules. THAT PATH IS NOT YET IN SERVICE. Enabling it depends on a key-custody arrangement that is being made outside this platform, and until that is settled the application keeps files on the application host instead. The description above is therefore the design, not a control that is currently operating, and it is written this way on purpose: a topology that describes intended controls as though they were live is worse than one that omits them. The service publishes an erasure guarantee: deleting a case destroys its content. The underlying store keeps object versions, so that guarantee is implemented by deleting every version of every object rather than by a single delete. The difference was verified rather than assumed: an ordinary delete makes an object unreadable while the content still exists, and a check that asserts only "the object is gone" passes in both cases. The compute migration is COMPLETE. The service previously ran on a single small host outside the managed estate; it now runs on the application tier, and that host's access to the production database has been withdrawn rather than merely unused. The remaining step is the object-storage path above. -------------------------------------------------------------------------------- PARTNER PORTAL https://partnerportal.rodmena.co.uk -------------------------------------------------------------------------------- Role Partner applications and partner records Database Region UK Depends on identity, mail_api Depended on by (none) DEPLOYED. This entry said NOT YET DEPLOYED until 2026-09-21, when the application replaced the placeholder page. The rest of the document was last verified on the date at the top; this block was checked on 2026-09-21. A company applies through a public form, confirms its address by a link sent to it, and a person reviews what arrives. Applicants who already hold an account sign in through identity. An application creates a record and starts a conversation; it grants no authority and forms no agreement, which the portal states on the form itself. The database is on the replicated tier with the same controls as every other: mutual TLS, a certificate whose subject must equal the role, and access restricted to the single address that serves the application. -------------------------------------------------------------------------------- HAVEN https://haven.rodmena.co.uk -------------------------------------------------------------------------------- Role Secure storage service Database Region FR Depends on auth Depended on by (none) -------------------------------------------------------------------------------- TEL https://tel.rodmena.co.uk -------------------------------------------------------------------------------- Role Telephony: outbound voice minutes and SMS, metered per tenant Database Region FR Depends on auth, tokengate Depended on by (none) -------------------------------------------------------------------------------- CONSENSUS https://consensus.rodmena.co.uk -------------------------------------------------------------------------------- Role Distributed strongly-consistent key-value store and coordination primitive (leader election, distributed locks, watches) Database None. It is a Raft cluster and holds its own data; it does not use PostgreSQL Depends on (nothing) Depended on by (available to any service; no first-party consumer yet) The cluster has three voting members in three regions (UK, FR, DE), with a Raft quorum of two, and survives the loss of any one member. Clients reach it through a single TLS edge, which currently runs in the UK region. Losing the edge's host makes the service unreachable even though the remaining members keep quorum, so the UK region is a single point of failure for clients. A second edge in a separate region is planned. This replaced a three-member cluster whose members all ran on one host. That arrangement survived a process restart, but it could not survive the host failure it was meant to guard against. The earlier documentation recorded that limitation. This was tested by killing the leader. The two surviving members elected a new leader and kept serving reads and writes. On rejoining, the recovered member reconciled the writes it had missed. Peer traffic between members is mutually authenticated with TLS against a private certificate authority. A Raft peer can commit entries, so an unauthenticated peer port would be a write path into the cluster. Client access uses TLS with per-user authentication and role-based authorisation. The cluster does not survive the simultaneous loss of two members. A single member has no quorum, so it stops accepting writes to avoid divergence. It then stops serving. No data is corrupted, and it recovers when a second member returns. ================================================================================ 3. DEPENDENCY GRAPH ================================================================================ Arrows point from a service to the services it calls at runtime. +------------------+ | AUTH | (no dependencies) +------------------+ ^ +----------+-----------+-----------+----------+ | | | | | +-------+ +--------+ +--------+ +-------+ +------+ | FUTEX | | MAIL | |AGENTBUS| | RED9 | |RUNFLOW| +-------+ | API | +--------+ +-------+ +------+ | +--------+ | | | | ^ ^ | | | | | | | | | +----------+ +----------+----------+----------+ | +--------------+ | TOKENGATE | (no dependencies) +--------------+ ^ | +--------------+ +-----------+ | IDENTITY |<---------| RED9 | +--------------+ +-----------+ Read in dependency order, lowest first: tier 0 auth, tokengate depend on nothing tier 1 identity, runflow depend on tier 0 (+ mail_api) tier 2 mail_api, futex mutually reference runflow tier 3 agentbus, red9, ledger consume the tiers below Practical consequences: - auth unavailable -> authorisation fails platform-wide - mail_api unavailable -> no outbound email; six services degrade - tokengate unavailable -> quota decisions degrade; authorisation is unaffected - agentbus unavailable -> inter-agent messaging stops; user-facing services are unaffected ================================================================================ 4. DATA TIER ================================================================================ PostgreSQL 18.6 on FreeBSD 15, one dedicated host per group of databases. Every cluster has data-page checksums on and uses the "builtin" locale provider with a C.UTF-8 collation, which is immune to libc collation drift across upgrades. HOST GROUP REGION PLATFORM DATABASES ---------------- -------------------- ----------------------------- Primary A United Kingdom futex, ledger, runflow, knowledge_base, trace, vellum, speakup, requests Primary B France auth, tokengate, identity, tel, provenance, haven Primary C United Kingdom agentbus, red9, mail_api, ci, tracker, partnerportal Standby Germany live replicas of every database on A, B and C Total data volume is small, and capacity is not a constraint at any tier. The largest single database is under 8 GB and the whole fleet is about 20 GB. The design constraints are isolation and recoverability. (Vellum moved from Primary B to Primary A on 2026-09-21 for UK residency.) Primaries A and C are in the same UK facility, 0.59 ms apart. This is a known concentration -- one facility incident reaches two thirds of the platform's primaries -- and the German replica set described below mitigates it. The two hosts are not independent of each other. The round-trip time was measured: two hosts are not treated as independent failure domains until someone has checked. ================================================================================ 5. ACCESS CONTROL AT THE DATABASE LAYER ================================================================================ No database on this platform accepts a password alone from the network. Every remote connection must present three independent factors: 1. TLS 1.2 minimum (TLS 1.3 with AES-256-GCM in practice), verified against a publicly trusted server certificate 2. A client certificate issued by a private certificate authority, whose subject common name must equal the database role being used 3. A SCRAM-SHA-256 password A connection lacking any one of these is refused before authentication completes. Unencrypted connections from the network are rejected. Isolation is enforced by host-based access rules as well as by grants. No service's role has a rule permitting another service's database, so a cross-database connection is refused at connect time, even with valid credentials. This is checked continuously. The certificate authority's private key is held offline. It is not stored on any database host, so a compromised database server cannot mint new identities. Two services, red9 and ledger, also enforce row-level security inside the database for per-tenant isolation. Their application roles are explicitly configured so that they cannot bypass those policies. ================================================================================ 6. REPLICATION ================================================================================ Every primary cluster streams to a physical standby in Germany, giving a continuously current copy of every database on the managed fleet outside both the UK and French facilities. Primary A (UK) --\ Primary B (FR) ----> Germany : 3 standby clusters, one per primary Primary C (UK) --/ Properties: - Replication is physical (block-level). It is exact, and it cannot silently diverge when a schema changes. - Replication is asynchronous, so a replica cannot slow down or block a primary. - No replication slots are used. A standby that falls behind recovers from archived write-ahead logs in object storage, so its primary does not have to retain them and a lagging replica cannot exhaust a primary's disk. - Replicas are read-only and are configured never to influence maintenance behaviour on their primary. - Promotion is a manual operation. The database tier has no automatic failover and no consensus layer. This is a considered trade-off: at this data volume, split-brain and consensus faults are harder to diagnose and recover from than the hardware failure they would guard against. (The consensus service in section 2 is a separate product with its own Raft cluster. It takes no part in database failover, and no database depends on it.) Replica health is checked automatically every ten minutes. The check fails on: instance down, unexpected promotion, replication not connected, replica connected but not streaming, or staleness beyond a threshold. The check was tested by severing replication and confirming that it reported the fault while the replica was still running and answering queries. IMPORTANT: replication is not a backup. A destructive or erroneous statement reaches every replica within milliseconds. Recovery from a logical error uses the backup system described below. ================================================================================ 7. BACKUP AND RECOVERY ================================================================================ Every primary cluster in the managed fleet is backed up continuously to off-site object storage in a separate facility, independent of any database host. Full backup daily Incremental backup hourly Write-ahead log archived continuously (at most every 5 minutes) Recovery point (RPO) 5 minutes or better, for the managed fleet Encryption AES-256-CBC, encrypted before leaving the host Retention 7 daily full backups, each with its hourly incrementals Point-in-time recovery is supported to any moment within retention, either in place or into a separate cluster, so a restore can be verified without disturbing production. Restore testing: the last physical restore test of the fleet (2026-08-09) wrote a known dataset, captured it in an incremental backup, restored it from object storage into a separate cluster, and found row count and content checksum identical. Separately, several services restore their nightly logical backups into a scratch database and require an exact row-count match before a backup is marked good. A full-size restore drill, to establish a measured recovery time, is planned; no recovery-time figure is published until it has been run. Every service also retains a pre-migration logical dump, held in two independent locations. Recovery objectives: Loss of one database host replica promotion, minutes Loss of an entire facility unaffected services continue; affected services recover from the German replica set or from object storage Logical error (bad statement) point-in-time recovery to just before it ================================================================================ 8. OPERATIONAL POSTURE ================================================================================ Database hosts - The firewall denies by default; only the ports required for operation are reachable. - Administrative access requires a public key; password authentication is disabled. - Repeated authentication failures trigger automated intrusion blocking. Blocks expire automatically, so no lockout can become permanent. - The base operating system is patched on a schedule, and security advisories are audited weekly. - Server certificates renew automatically, and the database reloads without operator action. Renewal is tested. Application hosts - The same posture: a default-deny firewall, public-key-only administrative access, automated intrusion blocking, unattended security updates, and a userspace out-of-memory guard configured to protect the supervision, container and administrative daemons rather than kill them. - Host metrics are collected from every host in the estate over TLS with authentication, not in clear text on a firewalled port. The firewall restriction is a control, not the only one. - A host that carries a workload but is not a monitoring target is treated as a defect rather than an omission. A missing target produces no alert, so "not monitored" and "no faults" look identical on a dashboard. - A host and the applications it serves have SEPARATE names, and the host's monitoring identity uses the host's own. Applications are retired, moved and renamed; a host whose metrics depend on an application's name loses them the day that application goes away, silently and at the worst moment. - Every host refuses TLS connections for names it does not serve, closing the connection rather than presenting some other name's certificate. A mismatched certificate is a warning users are trained to click through; a closed connection is not. Verification - A single re-runnable probe exercises every host: per-role connectivity, refusal without a client certificate, refusal of a wrong password, refusal of cross-database access, refusal of unencrypted connections, rejection of a certificate from an untrusted authority, negotiated TLS version and cipher, and remaining certificate lifetime. - Checks are designed so they can fail. Each check is validated against a known good case before it is relied on to report a bad one. Change management - Schema changes are applied by a migration tool whose state is stored inside the database it manages, so it survives restore intact. - Every service repository carries an operational runbook describing where its database lives, how to connect, how to run migrations, how to restore, and how to roll back. ================================================================================ 9. NOTES ON DISCLOSURE ================================================================================ The published version of this document omits: - host addresses and database endpoint names - database role and user names - object storage bucket names, regions and endpoints - internal service ports and filesystem paths - any credential, key, certificate or fingerprint None of these is needed to integrate with the platform's public APIs, and publishing them would materially help an attacker to target the infrastructure. Security issues may be reported to security@rodmena.co.uk. Reports affecting data integrity or isolation are treated as highest priority. ================================================================================ Rodmena Limited generated 2026-09-19 (UTC) ================================================================================