Cedar's fraud-agent asks to run a job on the Shared Runner
Four questions stand between the request and the work: who is asking, may it, does it fit right now, and can the destination check the answer? Read the picture left to right.
Statuses on this page: BUILT · RUNNING deployed and exercised end to end · BUILT · IN-PROCESS a tested library, not on the request path · DESIGNED · DRAFT on paper only · PARTIAL built, with a named gap. Each section heading carries its status once.
1What problem?
Autonomous agents from different customers share the same infrastructure. For every request we must prove who the agent is, decide what it may touch and by which route, reserve finite shared capacity without ever overspending, and let the destination check that the request really was granted that route.
The whole page uses one small cast. Remember these five names and the three numbers on the picture.
How this maps to the running system
| On this page | Canonical fixture id | In the deployed system |
|---|---|---|
| Cedar fraud-agent | fraud-detector (tenant cedar-bank) | the identity the test driver and six of the seven workload agents sign as |
| Northstar forecast-agent | demand-forecaster (tenant northstar-retail) | in the fixture and the Simulator; no deployed agent uses it |
| Fast Relay · 3 | fast-inference-relay, capacity 3 | entries in the catalog, not processes. The partitions count units on the links into and out of them. |
| Backup Relay · 2 | backup-inference-relay, capacity 2 | |
| Shared Runner · 4 | shared-inference-runner, capacity 4 | |
| transaction-features | risk-analytics-features | a file set served by the dataset-store, behind the sidecar |
The full fixture has five fictional customers: 48 graph nodes, 60 edges, 43 policy rules, 5 quotas and 25 requests. It also has two longer detours through regional relays. That is why the live policy service returns four routes for this request rather than two (§3). The relays and the runner are graph entries only. The only real destination behind the sidecar today is the dataset store.
2What model?
Three separate questions, answered by three separate parts. Keeping them apart gives clean failure reasons, and it stops one part from quietly doing another part's job.
Is this identity allowed this action on this resource, and which routes are legal?
Does every capacity-carrying hop on one legal route have a free unit at this moment?
The units are reserved, all together, under a lease that expires. A signed grant names that lease.
Authorized is necessary, not sufficient. A request can be allowed and still not fit. And nothing adaptive may ever turn "not allowed" into "allowed": feedback may only change preference among routes that policy already permits.
System map: what runs, and what is only designed
3How does a request work? BUILT · RUNNING
Follow the state card on the right. It shows what the system knows about the request after each step, and highlights what just changed.
0 · Ask
The agent signs a short assertion with its own key: RS256 only, valid for at most 60 seconds, with a unique id. It sends the assertion with a request id, the target (the Shared Runner) and the action (execute).
- identity
- —
- permission
- —
- eligible routes
- —
- chosen route
- —
- capacity
- —
- grant
- —
1 · WHO? AuthN fixes the identity
authn checks the signature, the lifetime and the algorithm, and refuses tricks such as alg=none. It remembers every (agent, assertion id) pair, so sending the same assertion twice gets REPLAYED. The agent's tenant and environment come from its own catalog row, never from the request.
- identity
- fraud-agent · cedar-bank · prod
- permission
- —
- eligible routes
- —
- chosen route
- —
- capacity
- —
- grant
- —
2 · MAY? Policy returns the legal routes
policy evaluates the rules for this identity, target and action. If allowed, it lists up to 8 legal routes within 6 hops, cheapest first. For this request both direct routes cost 4, and the tie is broken by link id, so the Backup Relay route comes first. If denied, the request stops here and no capacity is touched.
- identity
- fraud-agent · cedar-bank · prod
- permission
- allowed: execute
- eligible routes
- via Backup Relay (4), via Fast Relay (4), 2 longer detours
- chosen route
- —
- capacity
- —
- grant
- —
3 · FIT? The ledger finds a route that fits now
authn hands the routes to the manager, which picks one of two partitions (§6). The partition walks the routes in order and takes the first one on which every capacity-carrying link still has a free unit.
- identity
- fraud-agent · cedar-bank · prod
- permission
- allowed: execute
- eligible routes
- 4 routes
- chosen route
- via Backup Relay, the first that fits
- capacity
- —
- grant
- —
4 · COMMIT: the lease is reserved
In the same atomic ledger call, the partition reserves 1 unit on each capacity-carrying link of that route and returns a lease id. The lease lives between 5 and 30 seconds. Checking and reserving cannot be split by another request, so two requests can never both take the last unit.
- identity
- fraud-agent · cedar-bank · prod
- permission
- allowed: execute
- eligible routes
- 4 routes
- chosen route
- via Backup Relay
- capacity
- 1 unit per link, leased
- grant
- —
5 · GRANT: a signed path grant
authn signs a short-lived RS256 token that names the agent, the lease, the route, its exact hops and the policy version. Its expiry is capped below the lease's:
- identity
- fraud-agent · cedar-bank · prod
- permission
- allowed: execute
- eligible routes
- 4 routes
- chosen route
- via Backup Relay
- capacity
- 1 unit per link, leased
- grant
- signed; expires before the lease
6 · CHECK: the sidecar verifies
At the destination, the sidecar checks the signature against the published key set, then the audience, issuer and times, then that the hops it is shown are exactly the hops in the grant. One changed byte gives BAD_SIGNATURE. A different path gives PATH_MISMATCH.
When the driver calls the sidecar directly, the hops it shows are whatever the caller claims. The dataset store instead passes its own configured ingress chain (§10), so a grant for another route fails whatever the client sends. That chain is still configuration, not an observed network path; this is an open item (§12).
- identity
- fraud-agent · cedar-bank · prod
- permission
- allowed: execute
- eligible routes
- 4 routes
- chosen route
- via Backup Relay
- capacity
- 1 unit per link, leased
- grant
- verified by the sidecar ✓
Release, and what can go wrong
To give capacity back, the agent signs a fresh assertion and authn forwards release. Only the owner may release: another agent gets NOT_OWNER and the lease stays live. If nobody releases, the lease expires on its own; idle partitions sweep expired leases every second. Retrying the same request id returns the original answer instead of a second lease.
The eight recorded scenarios
A driver runs these against the deployed services. They use the read request from the cast: Cedar fraud-agent reads transaction-features through Cedar's gateway (two links, capacity 6 each). The runner request above takes the same code path but is not one of the eight.
| # | What happens | Expected and observed | Question it tests |
|---|---|---|---|
| S1 | allowed request | grant OK; sidecar OK | all four |
| S2 | a request policy denies | DENIED; partitions unchanged | may it? (no side effect) |
| S3 | admit until full | 6 admitted, the 7th FULL; nothing over capacity | does it fit? |
| S4 | release, then admit | admit OK | commit and release |
| S5 | tampered grant | BAD_SIGNATURE | check |
| S6 | different hops shown | PATH_MISMATCH | check |
| S7 | assertion sent twice | REPLAYED | who? |
| S8 | another agent releases | NOT_OWNER; lease still live | commit (ownership) |
All eight pass on a local three-node Minikube cluster and on a one-node managed cluster, including with the agent workload deployed (§9).
4Why graphs?
Each "may this identity do this action on this resource" rule decides whether an edge exists. Routes are then simply paths through the graph, and cross-tenant edges are absent by construction. That is the running part. Policy is compiled into a bitmap index, so one decision is a few AND operations. An independent "scan" evaluator that shares no code with the index must agree with it on every test case. Every allow must carry the owner-scope condition, and a deny that can never match is a compile error rather than a silent gap.
Max-flow: how much can run at once? BUILT · IN-PROCESS
Treat capacities as pipe widths. Max-flow asks: if every agent pushes as hard as it can, how many jobs reach the Shared Runner at the same time?
The policy gap: the real canonical result
Run max-flow twice, once on the physical network and once on only what policy allows, and the difference shows whether hardware or policy is the bottleneck. In the canonical fixture, Cedar's fraud-agent wants to read transaction-features at 4 units at once:
measured on canonical Computed by Dinic's algorithm, confirmed by an independent Edmonds–Karp oracle and by an exact assignment search (also 2). A quota-aware variant shows tenant shares binding: three runner requests of 2 units each (Cedar, Northstar and a third customer) fit 4 without quotas, but only 3 with each customer's declared share of the runner.
Max-flow is analysis and evidence, not live routing. The running system never computes a flow per request. The policy service lists legal routes and the partition takes the first one that fits (§3). Max-flow answers capacity-planning questions and checks that the limits are what we think they are.
Residual graphs: permission to undo BUILT · IN-PROCESS
A residual graph marks how much room is left on every link. It also adds a reverse edge for every unit already sent. A reverse edge is not traffic flowing backwards. It is permission to undo an earlier allocation so that the total can grow.
This rearranging happens inside the max-flow and assignment algorithms. The running system never moves a granted lease. Moving capacity between partitions is a separate, designed-only idea (§6).
5Why a ledger? BUILT · RUNNING
Planners may suggest; only the ledger commits. A request is a demand vector: one entry for every limited thing on its route. The vector either fits completely and is reserved in one step, or nothing is reserved.
How this maps to the running system
The live ledger counts links, not boxes. A lease takes exactly 1 unit on every capacity-carrying link of its route. For the Backup Relay route that means three links: Cedar's own link into its gateway (capacity 6), the link into the Backup Relay (2) and the link from it to the Shared Runner (2). The runner's overall limit of 4 and the customers' quota shares are used by the analysis in §4. The running partitions do not reserve them yet PARTIAL.
The ledger takes the current time as an argument, so every expiry boundary can be tested without sleeping, and it returns a named reason for every outcome. In the deployed system, S3 fills a route until it is full: the 7th admit is FULL and no link is ever above capacity. S4 shows that a release frees capacity immediately. S8 shows that only the owner can release.
The grant never outlives its lease. Its expiry is capped below the lease's (§3, step 5). One gap is disclosed: there is no revoke. A released lease frees capacity immediately, but a grant issued for it stays valid until it expires, at most 60 seconds later.
6Why partitions? BUILT · RUNNING
One ledger is one point of contention. So capacity is split across two partitions, and each partition owns a disjoint share of every limit outright. No two processes can ever both believe they own the same unit, so "never overspend" stays a simple, local property of each partition.
- Choice: "power of two choices". With two partitions the manager compares both partitions' last published spare room. Ties are broken by a hash of the request key, so the same request always lands in the same place.
- One fallback: on a definite
FULLthe manager tries the other partition once. That is why S3 admits all 6 units on a capacity-6 route (3 + 3) before the firstFULL. - Retries: re-sending the same request key returns the original answer from the partition that holds it, never a second lease.
The price of ownership is fragmentation. A route can fit globally yet fit in neither partition, for example if one link has room only in partition 0 and the next only in partition 1. The system does not measure how often this happens yet. Scaling is also unmeasured: the partition claim is functional only PARTIAL. Moving capacity between partitions is designed, not built. Fencing a restarted or duplicate manager is built (§10).
7Why fairness? DESIGNED · DRAFT
Policy decides who may compete. Capacity decides what fits. Neither decides who goes next when Cedar and Northstar both have work waiting. First-come-first-served lets a flood from one tenant crowd out the other. Deficit round robin (DRR) gives each tenant a fixed allowance of credit per round, and an expensive job spends more of it.
Allowance 2 per round. Cedar's job costs 3, Northstar's costs 1. Round 1: Cedar 0 + 2 = 2 (3 > 2, waits) Northstar 0 + 2 = 2 (1 ≤ 2, runs; 1 left) Round 2: Cedar 2 + 2 = 4 (3 ≤ 4, runs; 1 left)
teaching numbers The running system has no queue: every lease is 1 unit per link, and a request either fits now or gets FULL.
Credit is a scheduling budget, not capacity. A tenant can have plenty of credit and still nowhere safe to run; capacity always belongs to the ledger. The design (per-tenant queues and batched admission inside a partition) is a reviewed draft with open findings. Nothing is built.
8Why PCA? BUILT · IN-PROCESS DESIGNED · DRAFT
Suppose the Fast Relay starts to behave differently: latency and errors begin to move together in a new way. No single number may cross a threshold, yet the shape of the measurements has changed. PCA finds the direction in which several measurements vary together. Drift shows up as that direction rotating.
Built (in-process): the maths. Timestamped samples become fixed windows, are scaled, and give a covariance matrix whose eigenvectors (Jacobi's method) are the directions. Windows are compared by the angles between their main directions. Every degenerate input returns a named status rather than a number: too few samples, a constant feature, near-tied eigenvalues, or a solver that did not converge. Results are checked against hand fixtures, power iteration, planted spectra and an independent distance computation.
Not built: the feedback. There is no telemetry pipeline. Deciding whether a drift is significant, turning it into a bounded penalty, and letting that penalty re-rank the legal routes are all designs. Whether PCA detects real drift usefully is untested. And by design, feedback may only change which allowed route is preferred. It can never grant a permission or create capacity.
9What is live today? BUILT · RUNNING
Seven core processes run in one Kubernetes namespace: authn, policy, the manager, partition-0 and partition-1, the sidecar and the read-only ui (this site). They exchange one-line JSON messages over TCP and share no code. Together they used about 230 MiB of memory. The same stack runs on a local three-node Minikube cluster and on a one-node managed cluster.
The agents and the data plane
A continuous workload adds a dataset-store and seven agents, for 15 pods in total. Each agent loops: sign, get a grant, download one dataset from the store, train a small model, release the lease, log one line. The store asks the sidecar to check every grant before it sends a single byte.
| Agent | Data · model | Expected |
|---|---|---|
agent-creditcard-fraud | credit-card fraud sample (10.6 MB, the largest) · softmax regression | OK; beats "never fraud" |
agent-credit-default | credit-card default (2.9 MB) · softmax regression | OK; beats the majority class |
agent-german-credit | German credit (50 KB, fast loops) · softmax regression | OK; beats the majority class |
agent-mnist | handwritten digits (1.4 MB) · a tiny CNN, the heaviest training | OK; far above 10% chance |
agent-wine | red wine quality (0.1 MB) · small neural network, regression | OK; error below "guess the average" |
agent-titanic | Titanic passengers (61 KB) · softmax regression | OK; mostly measures the loop itself |
agent-denied-fleet | none | DENIED every time; no lease, no data |
The six allowed agents share one identity and one route with 6 units per link (3 per partition), so they compete for the same capacity. The seventh agent's tenant does allow that read, but an environment-wide deny for dev overrides the allow. Measured results are on the Journey page.
Network policy and the public UI
- Default-deny network policy. Each pod may receive connections only from named peers: only
authnand theuimay reach the manager, only the manager may reach the partitions, and only the six allowed agents may reach the store. It is enforced on the managed cluster (11 of 11 probes behaved as expected). The local cluster's network plugin ignores network policies. - Public UI. This read-only site is also published over TLS only, through one cloud load balancer with a connection throttle. It serves synthetic data only. Every other service is internal.
- Logs. Every process writes one JSON line per event with a trace id, so one request can be followed across processes. Keys and token values are never logged.
Claims hold for clients that use the front door. Traffic between services is plaintext and mutual TLS is deferred. The network policy narrows who can try, on the managed cluster only.
10How is it deployed? BUILT · RUNNING
This chapter follows the running system from the outside in: which processes exist and what each one owns, which Kubernetes objects hold them, who may talk to whom, and how a request and its data move through. Everything here was read from the renderer's output, the code and the live namespace (read-only), not from memory.
The components and what each owns
| Process | Code | Owns | Status |
|---|---|---|---|
policy | poc/adaptive-flow | The catalog and the rules. Answers may this agent do this? and, if yes, lists up to 8 legal routes, cheapest first. Holds no capacity. | RUNNING |
manager | poc/atomic-lease | Routing only: which partition to try first, and a memory of where each request key went. Owns no units itself. | RUNNING |
partition-0, partition-1 | poc/atomic-lease | Each owns a disjoint share of every link's units, the live leases on that share, and the replay records for 60 seconds. | RUNNING |
authn | poc/authn-sidecar | The issuer private key and the agents' public keys. Checks who is asking, calls policy and the manager, signs the path grant. | RUNNING |
sidecar | poc/authn-sidecar | Only the public key set. Verifies a grant; it cannot sign one (its code never imports the signer). | RUNNING |
ui | poc/adaptive-flow | This site and the read-only /api/snapshot. The TLS certificate. No write path of any kind. | RUNNING |
dataset-store | scripts/poc | The seeded dataset files and their manifest. Serves a file only after the sidecar accepts the grant. | RUNNING |
| 7 agents | scripts/poc | Each holds one agent private key and loops: sign, grant, download, train, release (§9). | RUNNING |
loadgen Job | scripts/poc | Only during a capacity run: one pod that drives the same grant and download path at stepped load, signing as the allowed agent. Deleted an hour after it finishes. | RUNS ON DEMAND |
The calls only go one way: agents → authn → policy, then manager → partitions; and for data, agent → dataset-store → sidecar. The ui reads from the manager, policy and the store's stats port. Nothing calls back into an agent. The members share no code: every hop is one-line JSON over TCP (NDJSON), or HTTP for the store and the ui.
Partition-owned leases, with epoch fencing BUILT · RUNNING
A lease lives only in the partition that granted it. The manager never holds units, so a lost manager loses no capacity. To stop two managers from both writing (for example an old pod stuck while a new one starts), every partition has exactly one owner at a time:
- Owner pair. The owner is the pair
(epoch, manager_id). The id is 16 random hex characters chosen at process start; the epoch is a number that only grows. - Takeover. A starting manager reads every partition, picks an epoch above any it saw, and configures each partition with its own pair. Leases and replay records survive the takeover. Then it re-reads all partitions; only when every one shows its pair does it start serving.
- STALE_EPOCH. Every
admitandlookupcarries the pair. A partition refuses any other pair withSTALE_EPOCHand writes nothing.releaseis not fenced: it can only free units, and it is already checked against the lease owner. - Superseded is final. A manager that sees a greater pair than its own stops serving for good and reports not ready. Only a process restart brings it back, as a new manager with a new id.
Spec docs/superpowers/specs/2026-10-02-partition-epoch-fencing.md; code poc/atomic-lease/partition.py (configure, admit, lookup) and manager_server.py (_on_foreign, _poll_round, skip_lookup). The code in the cluster's ConfigMaps was compared with the working tree on 2026-10-02 and matched file for file.
What the renderer creates
One script, scripts/poc/render.py lke, writes every object (40 of them) as JSON files. It refuses to write anything if its own checks fail (see the gates below). Then scripts/poc/deploy.sh applies them.
| Kind | Objects | Notes |
|---|---|---|
| Namespace | authnz-poc | Everything below lives in it. |
| ResourceQuota | authnz-poc-quota | CPU requests 1500m, memory requests 1700Mi, memory limits 2600Mi; at most 1 LoadBalancer Service, 1 NodePort (the one the load balancer needs; 0 on Minikube), 0 volume claims. |
| ConfigMap | code-adaptive-flow, web-adaptive-flow, code-atomic-lease, code-authn-sidecar, code-poc-scripts | The source code and this site. An init container copies them into a scratch volume at pod start. There is no custom image and no registry: every pod runs the official python:3.14-slim image, pinned by digest. |
| ConfigMap (by deploy.sh) | jwks | The public key set that authn publishes and the sidecar trusts. |
| Service | 8 ClusterIP + 1 LoadBalancer | See the port table below. |
| Deployment | 15 | One replica each, Recreate strategy (never two pods of one name at once). |
| NetworkPolicy | 9 | A default deny plus one allow rule per receiver (diagram below). |
| Job | loadgen | Only during a capacity run, rendered by scripts/poc/loadgen_job.py with the same checks. One pod, no retries, a hard deadline. |
Services and ports
| Service | Type | Port | Serves |
|---|---|---|---|
policy | ClusterIP | 7101 | NDJSON: decide, catalog, health |
manager | ClusterIP | 7201 | NDJSON: admit, release, snapshot, health |
partition-0 / partition-1 | ClusterIP | 7211 / 7212 | NDJSON: configure, admit, lookup, release, snapshot, health |
authn | ClusterIP | 7301 | NDJSON: grant, release, jwks, health |
sidecar | ClusterIP | 7302 | NDJSON: verify, health |
dataset-store | ClusterIP | 8090, 8091 | HTTP. 8090: /datasets/…, /health, /stats. 8091: /stats only, everything else 404. |
ui | ClusterIP | 8080 | HTTP for kubectl port-forward; no in-cluster allow rule reaches it. |
ui-public | LoadBalancer | 443 → 8443 | The only public entry: TLS straight to the ui pod. A NodeBalancer with a client connection throttle of 20. |
Pods: requests and limits
| Deployment | CPU request / limit | Memory request / limit | Secret mounted |
|---|---|---|---|
policy, manager, partition-0, partition-1 | 50m / 500m | 64Mi / 160Mi | none |
authn | 50m / 500m | 64Mi / 192Mi | authn-keys |
sidecar | 50m / 500m | 64Mi / 192Mi | none (public keys only) |
ui | 50m / 500m | 64Mi / 128Mi | poc-tls |
dataset-store | 100m / 500m | 48Mi / 128Mi | none |
agent-creditcard-fraud, agent-mnist | 15m / 120m | 40Mi / 128Mi | agent-fraud-detector-key |
agent-credit-default | 15m / 100m | 40Mi / 96Mi | agent-fraud-detector-key |
agent-german-credit, agent-wine, agent-titanic | 10m / 80m | 32Mi / 64Mi | agent-fraud-detector-key |
agent-denied-fleet | 5m / 30m | 24Mi / 48Mi | agent-route-optimizer-key |
loadgen Job (capacity runs) | 250m / 1 | 128Mi / 256Mi | agent-fraud-detector-key |
Without the Job the set needs 530m CPU and 736Mi memory in requests, and 1888Mi in memory limits (the denied agent's 64Mi init container counts there). Source: the renderer's preflight_ok line.
Secrets
No Secret is ever rendered or kept in the repository. deploy.sh creates four from files outside the repository, through a pipe so no value is printed. The renderer's SECRET_OWNERS table fixes which workloads may mount each one, and the gate refuses any other reference.
authn-keys: the issuer signing key and the agents' public keys. Onlyauthn.agent-fraud-detector-key: the allowed agent's private key. Only the six allowed agents and theloadgenJob.agent-route-optimizer-key: the denied agent's private key. Onlyagent-denied-fleet.poc-tls: the public site's certificate and key. Onlyui.
All keys except the TLS certificate are synthetic test keys made by scripts/poc/keygen.py.
The same safety settings on every pod
- Runs as a non-root user (uid 10001), with no privilege escalation.
- Read-only root filesystem. Code and dependencies are written once by init containers, then mounted read-only.
- All Linux capabilities dropped. The seccomp profile is
RuntimeDefault. - No service-account token is mounted, and service links are off. No pod holds credentials for the Kubernetes API.
- Only three volume kinds: ConfigMap, Secret and size-limited scratch (
emptyDir). No persistent volumes. authn,sidecarand the agents need PyJWT. An init container installs it withpip --require-hashesfrom a hash-locked list, binary wheels only, into a read-only/depsmount.
What the gates refuse
render.py runs these checks before it writes, and deploy.sh runs them again (render.py --check) on the files on disk:
- More than one LoadBalancer Service; any NodePort Service or
nodePortfield; external IPs; source ranges anywhere but the public ui Service; a NodePort quota that differs from the LoadBalancer count. - Any volume claim, persistent volume or StorageClass; any object kind outside Namespace, ResourceQuota, ConfigMap, Service, NetworkPolicy, Deployment and Job (so no Secrets, RBAC or Ingress in the set).
- A replica count above 1; a Job with parallelism above 1, retries, no deadline, or a deadline or cleanup time above two hours.
- Any pod that breaks the safety settings above, uses another image, or sets a host port.
- A Secret mounted by a workload not in
SECRET_OWNERS; the NetworkPolicy labels on any workload other than their intended ones. - A NetworkPolicy set that is not exactly the expected one, or that lacks the default deny; a set whose requests exceed the quota; a ConfigMap of 900 KiB or more.
The network
kubectl port-forward does not pass through these rules, which is how the driver and operators reach pods. Sources: scripts/poc/render.py (NETWORK_POLICIES), docs/operations/lke-agents.md (node firewall).| NetworkPolicy | Receiver | Allowed senders | Port |
|---|---|---|---|
default-deny-ingress | every pod | nobody | any |
allow-policy | policy | authn, manager, ui | 7101 |
allow-manager | manager | authn, ui | 7201 |
allow-partitions | partition-0, partition-1 | manager | 7211, 7212 |
allow-authn | authn | pods labelled authnz-poc/role=agent | 7301 |
allow-sidecar | sidecar | authn, dataset-store | 7302 |
allow-dataset-store | dataset-store | pods labelled authnz-poc/dataset-access=allowed | 8090 |
allow-dataset-store-stats | dataset-store | ui | 8091 |
allow-ui-public | ui | any address (via the load balancer) | 8443 |
The gate also checks that the two agent labels sit only on the workloads meant to carry them, so a new pod cannot give itself store access by copying a label.
The edge of the cluster
- One load balancer. The
ui-publicService makes the cloud create one NodeBalancer. It passes TCP 443 straight to the ui pod's port 8443; TLS ends in the ui process, using the site certificate. A connection throttle of 20 is set by annotation. - Node firewall. A Cloud Firewall sits in front of the node: inbound traffic is dropped by default. It admits only the cluster's own control traffic from the provider's private ranges, and NodePorts only from the NodeBalancer's backend range, so the load balancer's NodePort is no longer a public side door. No public SSH. Outbound is open. Akamai's Cloud Firewall Controller keeps it attached, and
scripts/poc/firewall_reconcile.shre-checks it after every deploy.
The firewall rules are taken from the runbook (docs/operations/lke-agents.md); no probe of the node's public NodePort is recorded since the firewall was added.
The cluster
No persistent volumes, no autoscaling, no paid add-ons. A local three-node Minikube cluster runs the same manifests without the load balancer; its network plugin does not enforce NetworkPolicy.
How data moves
Grant and download, one iteration
poc/authn-sidecar/authn_server.py (Authn.grant), poc/atomic-lease/manager_server.py (_admit_locked), scripts/poc/dataset_store.py, scripts/poc/agent_worker.py (Worker.iterate).- Ask. The agent signs a fresh RS256 assertion (30-second life, unique id) and sends it with a new request id, the target and the action.
- Who.
authnchecks the signature with that agent's public key, the lifetime (at most 60 seconds, 5 seconds leeway) and the claims. Its anti-replay cache remembers each (agent, assertion id) until the assertion expires; a repeat isREPLAYED. When the cache is full it refuses rather than forgets. - May.
policydecides from the agent's own catalog row and returns up to 8 routes. - Route. The
managersends a key it remembers straight to its partition. A new key is first looked up on both partitions, unless the lookup can be skipped: the manager holds every partition, its key memory has never dropped an entry, and no partition holds records from an earlier owner. Then it tries the partition that power-of-two-choices picks: two candidates from the key's hash, the one whose routes have more spare units first. - Fit. The partition checks the owner pair, then reserves 1 unit on each capacity-carrying link of the first route that fits, all or nothing. If it is
FULL(or restarted and unconfigured), the manager tries the other partition once. Any other answer goes straight back. - Grant.
authnsigns a path grant naming the lease, route and hops; it expires before the lease does. - Download. The agent sends the grant to the
dataset-store. The store asks thesidecarover a reused connection, passing its own configured ingress chain as the hops, never the client's claim. Only onOKdoes it stream the file, and only a file whose size and sha256 matched the manifest. The kernel copies it (sendfile), with a plain read loop as fallback. - Release. After training, the agent signs a new assertion and releases the lease through
authnand the manager. If nobody releases, the lease expires by itself (5 to 30 seconds).
The ui snapshot RUNNING
When a browser asks for /api/snapshot, the ui refreshes at most once a second: it asks the manager for its snapshot (which includes every partition's), policy for the catalog, and the store's port 8091 for /stats. Each call has a 0.8-second budget. If the store or policy fails, the view still comes back, with that part marked. If the manager fails, the last view is served, marked cached, for up to 5 seconds, then the answer is an error. The view goes to the browser over TLS, through the load balancer.
Seeding the store BUILT
The store downloads nothing. An operator runs scripts/poc/seed_datasets.sh on a laptop: it checks the local files against their manifest, copies each into the store pod with kubectl cp under a temporary name, then renames them all into place in one step, manifest last. The store's catalog re-reads every 5 seconds and serves a file only after its own size and hash check, so it never sees half a file. The script waits until the store reports this exact manifest. The data sits in a scratch volume and is lost when the pod restarts.
Logs and evidence BUILT
Every process writes one JSON line per event to standard output, with a trace id that follows a request across processes. Key and token values are never logged. After a run, the driver saves each pod's log, and scripts/poc/collect_logs.py merges them into one time-sorted logs.jsonl (non-JSON lines go to a side file). Results stay in a local, untracked poc-results/ folder.
11How is the code built? BUILT · RUNNING
Inside each process: which modules import which, the main classes and what each one guarantees, how main() wires them up, the data each process keeps and its limits, and the order in which each request is checked. Every fact cites its file and function. The code in the cluster matched the working tree file for file on 2026-10-02.
Rules every member follows
- Layers. Imports point one way: entry points and adapters → application → domain logic → domain data. Domain logic does no I/O and is given the time as an argument; sockets, clocks and environment variables live in the
*_server.pyentry points. - No shared code. No member imports another. Processes talk one-line JSON over TCP (NDJSON) or HTTP, so each member has its own small NDJSON module.
- Standard library only, except where named:
poc/authn-sidecaruses PyJWT withcryptography(pinned and hash-locked, a user-approved exception), and the agent pods borrow that same install for signing.
| Code | Import-graph test | What it enforces |
|---|---|---|
poc/atomic-lease | tests/test_multi_layering.py ENFORCED | An exact list of allowed imports per file (read from the syntax tree, including imports inside functions); every file declared; no cycles. The ledger module may not use await, try or the clock, so each ledger call is atomic. |
poc/authn-sidecar | tests/test_layering.py ENFORCED | An allowed import graph; the pure modules (issuer, verifier) import no I/O module; the sidecar side can never import the signing code. |
poc/adaptive-flow | tests/test_layering.py ENFORCED | Every file has a declared layer 0 to 3; imports only to the same or a lower layer; application never imports an adapter; no cycles; no dynamic imports. Project imports only, not third-party ones. |
scripts/poc | none DESIGN RULE | The store and agent follow the same shape (pure agent_models, I/O in the entry points), but no test checks it. |
Partitions and manager · poc/atomic-lease
lease_config.py (reads TTL settings from pyproject.toml) and the legacy ledger.py are not used by the servers. Source: tests/test_multi_layering.py (DECLARED).Classes and what each guarantees
| Class | Holds | Guarantee |
|---|---|---|
Partitionpartition.py | Owner pair, capacities, the ledger, records by key and by lease id, expiry heaps, counters. | Refuses everything (FENCED) until configured. Every call fully applies or changes only the reject counters. Its clock never moves back. One writer. |
_Recordpartition.py | Owner, request key, the routes offered, the stored reply, replay horizon, lease expiry, epoch, ended flag. | Written only for a successful admit, so a refused request can simply be retried. |
MultiLedger, TtlBoundsmulti_ledger.py | Units per link, live leases, spent ids. | A lease takes one unit on every link of its route or nothing. TTL bounds satisfy min ≤ max ≤ replay window. |
PartitionServicepartition_server.py | The partition, its clocks, a 1-second timer. | Expires due leases before every operation and once a second when idle. |
Managermanager.py | Remembered key → partition (an ordered dict), the last known spare units, outcome counters. | Pure: no I/O, no clock. Once any key is dropped from memory, a flag stays set for the life of the process. |
ManagerServicemanager_server.py | Its manager_id, per-partition epoch, owned and stale-free flags, the confirmed flag, the superseded flag, per-key locks. | Serves admits only while it holds every partition: owned, confirmed by a snapshot round taken after the last configure, and not superseded. |
Wiring and concurrency
- Partition (
partition_server.main): flags--host,--port; environmentPROC_NAME,MIN_TTL_S(5),MAX_TTL_S(30),REPLAY_TTL_S(60),MAX_TRACKED_IDS(65,536). It buildsTtlBounds→Partition→PartitionService→ an NDJSON server. One asyncio event loop; every partition call is synchronous, so no locks are needed. - Manager (
manager_server.main):POLICY_ADDR,PARTITION_ADDRS,PROC_NAME, optionalEPOCH(else the current time in seconds). It buildsManagerService(which buildsManager) and the server, then a background task fetches the catalog from policy (retrying every second) and starts the poll loop: every second, snapshot the partitions, configure any it does not own, and run the confirm round. One event loop, oneasyncio.Lockper request key (dropped when nobody waits), at most one configure per partition at a time. - NDJSON (
ndjson.py): lines up to 64 KiB, 128 connections (more getBUSY), 10-second read and write timeouts, 30 seconds per request. The client opens one connection per call with a 5-second timeout.
Data and limits
| Data | Shape | Bound |
|---|---|---|
| Lease | id pid:epoch:n; n counts successful admits | life clamped to 5–30 s |
| Replay record | by request key and lease id | kept until admit + 60 s; at most 65,536 tracked ids (TRACKING_FULL) |
| Owner pair | (epoch, manager_id), compared in that order | epochs only grow; a configure never reuses one |
stale_records | count of records from an earlier owner | set at takeover, falls as they age out (60 s) |
| Routed-key memory | ordered dict, most recent last | 65,536 keys; oldest dropped first, which turns the lookup skip off for good |
| Requests | at most 8 routes of at most 64 hops; ids up to 256 characters | larger is INVALID |
Partition admit, in order (partition.py · Partition.admit)
| # | Check | Answer |
|---|---|---|
| 1 | Not configured yet | FENCED |
| 2 | Bad key, owner, routes, TTL, wall time or pair | INVALID |
| 3 | Pair is not the owner pair | STALE_EPOCH, with the current pair; nothing written |
| 4 | Clock not finite or moved back | CLOCK |
| 5 | A record exists for this key: different owner or routes / ended or expired / otherwise | CONFLICT / ENDED / the stored answer, replayed: true |
| 6 | Try routes in order: one unit per capacity-carrying link, all or nothing; a full route moves on | OK with lease id, route, hops and expiry; or TRACKING_FULL |
| 7 | No route fits | FULL |
lookup runs checks 1–5 and answers NOT_FOUND instead of admitting. configure: first time → OK; other pid or capacities → CONFLICT; same pair → OK; a greater epoch → OK, bumped, leases kept; anything else → STALE_EPOCH. release checks the lease owner (NOT_OWNER) but not the pair.
Manager admit, in order (manager_server.py · _admit, _admit_locked, _lookup)
| # | Condition | Then |
|---|---|---|
| 1 | Bad fields / no routes / more than 8 | INVALID / NO_FEASIBLE / INVALID |
| 2 | No catalog yet, or not holding every partition | UPSTREAM_UNAVAILABLE, nothing sent |
| 3 | Key remembered | go to 5 with its partition first |
| 4 | New key: skip the lookup only if (a) holding and not superseded, (b) memory never dropped a key and every sent admit is remembered, (c) every partition reported stale_records = 0 under our pair. Otherwise look up on both partitions at once. | found → return that answer; STALE_EPOCH or an error → UPSTREAM_UNAVAILABLE; not found → 5 |
| 5 | Order: hash the key to two candidates; try first the one whose best route has more spare units (power of two choices) | send admit with our pair |
| 6 | Answer OK | remember the partition; return the lease |
| 7 | Answer FULL or FENCED (a definite no-write), or the connection failed | try the other partition once (a failed connection on a remembered key's own partition is UPSTREAM_UNAVAILABLE instead) |
| 8 | STALE_EPOCH, or an error after sending | UPSTREAM_UNAVAILABLE, no fallback (the admit may have landed, so the key is remembered) |
| 9 | Any other answer | returned as it is |
| 10 | Both tried | FULL if either was full, else UPSTREAM_UNAVAILABLE |
A foreign pair seen anywhere (_on_foreign): smaller than ours, or our own id → take the partition back with a fresh epoch; greater → superseded, final. Release is not gated on holding and takes no key lock.
authn and sidecar · poc/authn-sidecar
tests/test_layering.py checks this graph, that issuer and verifier import no I/O module, and that the sidecar side never reaches issuer.Classes
| Class | Guarantee |
|---|---|
JtiCache issuer.py | Remembers each (agent, assertion id) until expiry + 5 s. At most 10,000 entries, 1,000 per agent. When full it refuses new assertions; it never evicts a live id. Owned by one event loop. |
Authn authn_server.py | Holds the issuer key, the agents' public keys, one JtiCache and a grant counter. At start, the published key set must match the private key, or it refuses to start. |
Config authn_server.py | Frozen settings, checked at start: lease TTL in [5, 60] s, upstream timeout in (0, 30] s. |
Rejected issuer.py, verifier.py | A reason the caller sees plus a detail that is only logged. |
Wiring
- authn (
authn_server.config_from,start):POLICY_ADDR,MANAGER_ADDR,ISSUER_KEY_PATH,AGENT_KEYS_DIR,JWKS_PATH,LEASE_TTL_S(30),UPSTREAM_TIMEOUT_S(5) → keys loaded →Authn→ NDJSON server on one asyncio loop. Upstream calls open one connection each. - sidecar (
sidecar_server): onlyJWKS_PATH; the key set is read once at start. - Both: 64 KiB lines, 128 connections (extra ones are closed unanswered), 10-second timeouts.
The grant and its checks
A path grant carries iss, sub, aud = sidecar, nbf, exp, jti (lease id plus a counter), lease_id, route_id, hops, policy_version, target and action, signed RS256 with key id issuer-1. Its expiry is min(now + 60, lease expiry − 7), rounded down (issuer.grant_exp).
| Step | authn grant (Authn.grant, issuer.verify_assertion) | Answer |
|---|---|---|
| 1 | Fields missing or longer than 128 characters | MALFORMED, connection closed |
| 2 | Token over 8 KiB or unreadable; algorithm not RS256; a key-fetching header (jku, x5u, jwk, x5c, crit); no key id; unknown key; bad signature, audience or issuer; lifetime not in (0, 60] s; issued in the future; expired | BAD_ASSERTION (the detail is only logged) |
| 3 | Assertion id already seen / cache full | REPLAYED / BAD_ASSERTION |
| 4 | Policy unreachable / not allowed / no routes | UPSTREAM_UNAVAILABLE / DENIED / NO_ROUTE |
| 5 | Manager unreachable / refuses | UPSTREAM_UNAVAILABLE / its reason, e.g. FULL |
| 6 | The manager replayed an earlier admit | REPLAYED_REQUEST; no second grant |
| 7 | Hops not among the offered routes / lease too short to sign | release, then UPSTREAM_UNAVAILABLE / LEASE_TOO_SHORT |
| 8 | Otherwise | OK with the grant |
| Step | sidecar verify (sidecar_server.verify, verifier.verify_grant) | Answer |
|---|---|---|
| 1 | Request malformed (more than 64 hops, long fields) | MALFORMED |
| 2 | Token unreadable, not RS256, key-fetching header, unknown key id, bad signature | BAD_SIGNATURE |
| 3 | Wrong audience or issuer, missing claim, wrong types, target or action differs, not yet valid | BAD_CLAIMS |
| 4 | Expired (5 s leeway) | EXPIRED |
| 5 | Hops shown ≠ hops signed, as an exact list | PATH_MISMATCH |
| 6 | Otherwise | OK with agent, lease, route and hops |
policy and ui · poc/adaptive-flow
serve_policy also imports two error types from policy.bitmap and algo.paths; routes also imports graph.types. The ui imports nothing of the policy code: it reads policy over the network. Source: tests/test_layering.py (DECLARED_LAYERS).Classes
| Class | Guarantee |
|---|---|
PolicyIndex, Decision policy/bitmap.py | Rule i is bit i. Compiling refuses unknown conditions, rules that can never match, and allow rules without the tenant-scope condition. Single writer. |
Service, Query routes.py | Frozen after prepare: the graph, the policy index, the hop limit, edge costs, and each agent's tenant and environment. |
State serve_policy.py | Frozen: the catalog and the service, built once from the seed bytes (their sha256 is the catalog version). |
Server, Handler serve/httpd.py | A thread per connection under a global slot limit and a per-address limit. One request per connection; only GET and HEAD. |
App, SnapshotCache serve/httpd.py, serve/views.py | The only shared mutable state, under a lock. Refreshes at most once a second, by one thread at a time; never serves a view older than 5 s as a success. |
Wiring
- policy (
serve_policy.main):HOST,PORT,SEED_PATH→ read the seed once →routes.catalog,routes.prepare→ an asyncio NDJSON server (64 KiB lines, 128 connections, 10 s timeouts). A bad seed or policy stops it before it listens. It calls nobody. - ui (
serve.httpd.main):MANAGER_ADDR,POLICY_ADDR,STORE_ADDR,MAX_PER_IP(32 in the cluster),TLS_CERT_PATH,TLS_KEY_PATH,TLS_PORT. It loads the whole site into memory at start (only.html .css .js .json, no symlinks or dot-files, at most 1 MiB per file and 4 MiB in total), then runs a plain server and, with a certificate, a TLS server (TLS 1.2 or newer) on a second thread. At most 32 connections at once. - Every ui response carries a strict Content-Security-Policy (scripts from this site only, no frames) and no-sniff and no-referrer headers; TLS responses add HSTS.
Decisions
| Step | policy decide (routes.decide, PolicyIndex.evaluate) | Answer |
|---|---|---|
| 1 | Unknown agent, agent without an owner, unknown target | DENIED, no routes |
| 2 | Any matching deny rule (for example "environment is dev") | DENIED, even if an allow also matches |
| 3 | No matching allow | DENIED (default deny) |
| 4 | Allowed: search simple paths within the hop limit (6 for the fixture) and a budget of 100,000 search states | PATH_BUDGET if exceeded |
| 5 | Keep paths whose every link the agent's tenant may use; sort by cost, then link ids; keep 8 | OK with routes, or NO_ROUTE |
| Step | ui request (Server.process_request, read_head, views.normalize_path) | Answer |
|---|---|---|
| 1 | No free slot (global or for that address) | closed unanswered |
| 2 | Request line over 8 KiB / head over 16 KiB or 32 lines / not three words / 10 s head deadline | 414 / 431 / 400 / closed |
| 3 | A write method / an unknown method | 405 / 501 |
| 4 | Path with .., empty segments, backslash, control characters, bad UTF-8 / over 2048 bytes | 400 / 414 |
| 5 | /api/health, /api/snapshot, /api/seed, or a preloaded file | 200 (snapshot: 503 once the last good view is over 5 s old) |
| 6 | Anything else | 404; no request ever opens a file |
dataset-store and agents · scripts/poc
Classes
| Class | Guarantee |
|---|---|
Catalog, Entry dataset_store.py | A file is ready only if its size and sha256 matched the manifest; it is hashed again only when its size, modification time or manifest entry changes. Guarded by a lock. |
SidecarClient dataset_store.py | A pool of at most 16 idle connections, each used by one thread at a time; idle over 5 s means closed. A failure on a reused connection is retried once on a new one, but never after a timeout. |
Stats dataset_store.py | Counts by outcome and bytes served, under a lock. |
Worker agent_worker.py | Single-threaded; network, signing, randomness and clock are passed in, so tests can replace them. |
HttpResponse agent_worker.py | Reads the body in bounded lines or chunks and raises BadData if fewer bytes arrive than announced. |
Wiring
- dataset-store (
dataset_store.main):PORT8090,STATS_PORT8091,DATA_DIR,STORE_TARGETS,SIDECAR_ADDR,POLL_S5,INGRESS_HOPS(the two fixture links into the store). A thread per connection, at most 64 on the data port and 4 on the stats port (more get 503BUSY); one thread re-reads the catalog every 5 s. - agent (
agent_worker.main):AGENT_NAME,AGENT_ID,AGENT_KEY_PATH,AUTHN_ADDR,STORE_ADDR,TARGET,DATASET,EXPECT. One connection per call (10 s to authn, 30 s to the store). OnSIGTERMit releases its lease, then exits. - models (
agent_models.DATASETS): pure-Python softmax regression for four datasets, a one-hidden-layer network for wine, and a tiny convolutional network for the 14×14 digits. Each samples at most a fixed number of rows.
Decisions
| Step | store request (Handler.do_GET, read_head) | Answer |
|---|---|---|
| 1 | No free slot | 503 BUSY |
| 2 | Head limits (as the ui) / POST, PUT, DELETE, PATCH or HEAD | 414, 431, 400 / 405 |
| 3 | Stats port: anything but /stats | 404 |
| 4 | Bad path shape / bad target or file name / bad hop header | 404 / 400 BAD_PATH / 400 BAD_HOPS |
| 5 | No bearer grant, or over 8 KiB | 403 NO_GRANT |
| 6 | Client's hop claim differs from the store's ingress chain | 403 PATH_MISMATCH, before any verify |
| 7 | Sidecar unreachable / sidecar says no | 503 UPSTREAM_UNAVAILABLE / 403 with its reason |
| 8 | Not in the manifest / not seeded / size or hash wrong, or changed since checked | 404 NOT_FOUND / 503 NOT_SEEDED / 503 HASH_MISMATCH |
| 9 | Otherwise: stream with sendfile (read loop if unsupported) within 30 s + 1 s per 64 KiB | 200; stream outcome OK, SHORT_READ, TIMEOUT, STREAM_DEADLINE or CLIENT_ABORT |
| Stage | agent iteration (Worker.iterate, next_sleep) | Outcome |
|---|---|---|
| grant | Sign a 30 s assertion, ask authn: network error / refused | UPSTREAM_UNAVAILABLE / authn's reason, e.g. DENIED |
| download | GET the dataset with the grant: timeout / network error / not 200 | TIMEOUT / UPSTREAM_UNAVAILABLE / the store's reason |
| train | Sample rows, train, evaluate: bad data | BAD_DATA, else OK |
| always | Release the lease if one was granted; a failed release is logged and never changes the outcome | flag unexpected if outcome ≠ EXPECT |
| sleep | After UPSTREAM_UNAVAILABLE or TIMEOUT: min(30, 2n−1) s × a random 0.5–1; otherwise 2–5 s | then the next iteration |
12What remains open?
- The running ledger reserves link capacities only. Node limits (such as the runner's 4) and tenant quota shares are analysed, not reserved.
- The hop check compares the grant with a configured chain (the store's) or with the caller's claim (the driver's direct sidecar calls), never with the real path.
- Network policy is enforced on the managed cluster only, with no egress rules. The node firewall now limits the load balancer's node port to the load balancer, but no probe of that has been recorded yet.
- Policy grants are tenant-local only; sub-tenant inheritance is unsupported by decision.
- Partition scaling and fragmentation are unmeasured.
- No revoke: a grant can outlive a released lease by up to 60 seconds.
- PCA settings: tolerances of 1 or more are refused, but loose values below 1 (0.5, and 0.49 in a re-check) are still accepted and return "OK" with wrong eigenvalues.
- Fair queuing (DRR) and batched admission.
- Drift significance, the penalty state machine, and routing feedback.
- Rebalancing capacity between partitions and startup quarantine. (Epoch fencing of the partitions is now built: §10.)
- Mutual TLS between services.
- Grant renew and revoke.
- A telemetry pipeline.
The drafts still carry open review findings, which block the claims that depend on them. Their history is on the Journey page.
Glossary
| Permission | Whether policy allows an identity an action on a resource, and by which routes. |
| Feasibility | Whether a request fits every capacity limit on a route right now. |
| Commitment / lease | Capacity actually reserved, all together, until release or expiry. |
| Path grant | A short-lived signed token naming a lease, its route and its exact hops. |
| Residual | Room left on a link, plus reverse edges that allow an earlier allocation to be undone. |
| Partition | A ledger that owns a disjoint share of every limit. |
| Fragmentation | Capacity exists overall, but no single partition has the whole combination a route needs. |
| Oracle | An independent, deliberately simple implementation used to check the real one. |
Living document, checked against the code, the canonical fixture, the specs and the recorded runs. Read the Journey for how it was built and what changed on the way, and use the Simulator to explore the full fixture.