What was built in three days?
A working path from an agent's request to a check at the destination, running on Kubernetes:
- Authenticated agenta signed assertion, replay-proof
- Policy graphmay it, and by which legal routes?
- Partition-owned capacityall-or-nothing leases, never overspent
- Signed path grantnames the lease, the route and its hops
- Sidecar verificationat the destination, before any data moves
Here's the unusual part: most of the implementation and review work was done by coding agents operating under a rulebook. A human set the goals, made the consequential calls and held the keys to anything that costs money or leaves the machine.
How the system itself works is on the Docs page. This page is about how it was made, what changed on the way, and what evidence backs it.
1How did we build it?
A file at the root of the repository, the agent working agreement, binds every agent. It is policy, not magic: the project's living record states separately which rules are enforced by code and which are only followed. The agents run inside a harness that routes each job to a suitable model (strong models for hard work and review, light ones for searches and mechanical edits), loads written procedures (skills) when a task needs them, and checks that the living record still matches the code.
Say exactly how sure you are
Every claim carries a label: PASS, FAIL, PARTIAL, BLOCKED, INCONCLUSIVE, STALE or ASSUMED. "It works" is not allowed.
Why: when agents write most of the code, the danger is confident overstatement. A label makes "I didn't check that" visible, so a gap blocks only the claim that depends on it.
Write the failing test first
Watch it fail for the intended reason, then write the smallest code that passes. An import error does not count.
Why: a smoke detector has a test button. A test you have never seen fail may not be connected to anything.
Check every algorithm against an oracle
A second, deliberately simple implementation that shares no code: Dinic against Edmonds–Karp, bitmap policy against a plain scan, Jacobi against power iteration.
Why: to check a sum, add the column a second way. If both ways share a helper, a bug in it is "confirmed" twice.
Name the bug each test should catch
Plant a small deliberate bug (a mutant); the test must fail on it, for the right reason.
Why: a fire drill. Counting alarms proves nothing; lighting a match under one does.
Someone else reviews every change
A fresh, read-only reviewer sees the artifact, the requirements and the evidence, never the author's reasoning. One follow-up round; more need the human.
Why: you read what you meant, not what you wrote. Caveat: reviewers were mostly the same model family, so this is independence of context, not of vendor.
Unattended runs: one second opinion, hard limits stay human
A blocker gets written down with one recommendation and goes to one fresh advisor (AGREE or DISAGREE). Spending beyond a fixed cap, pushing code, DNS, real credentials and weakening a test are never the advisor's call.
Why: a house-sitter can call a plumber but cannot sell the house. One round stops two models talking each other into something.
Two code rules: the runtime uses the Python standard library only, except the member that signs tokens (two pinned, hash-locked cryptography packages that passed a supply-chain check); and imports flow one way through layers, enforced by tests in each member.
2Timeline
One line per phase. Open a line for the detail.
before 29 SepBaseline · a single-constraint lease ledger
A single-constraint lease ledger with 10 tests and one recorded 14-worker benchmark already existed. The repository history starts from it.
29–30 SepGraph algorithms · every algorithm against its own oracle
A 197-row canonical fixture with five fictional customers, then BFS, strongly connected components, topological order, Dijkstra, Dinic max-flow, bounded path enumeration and an exact assignment search. A census records the policy gap (4 → 2). Fitted scaling exponents came out at 0.98–1.00. The whole-branch review covered 40 commits and a 606-test suite: 0 critical, 0 important, 11 should-fix. Of 25 planted mutation probes, 23 were caught; new tests now catch the other 2.
30 SepPolicy bitmaps · a fast index that must agree with a plain scan
25 mutants, 0 survivors. Review found one important hole: an allow without the owner-scope condition fired for any tenant's agent. Such rules are now refused at compile time.
30 SepRules for unattended runs · the night loop, a spend cap, a dependency check
The rulebook gained its unattended-mode section, the standing cloud authorization with its spend cap, and a security check for any new dependency.
30 Sep – 1 OctMulti-constraint lease · all-or-nothing across many limits
Built beside the untouched original so the two can be compared. The lease suite grew from 10 to 72 tests, with 20 of 20 mutants killed.
night → 1 OctFirst unattended run · one controller per track
Finish the lease, then the PCA maths, then log evidence. Every blocker went through the one-advisor procedure (§5).
1 OctPCA maths · drift detection as a library
Windows, scaling, covariance, Jacobi eigenvectors, drift between windows and a whitened distance. Every degenerate case returns a named status. On the long benchmark Jacobi was 19× faster than power iteration at 4 features and 5.2× at 64.
1 OctLog evidence · results checkable from saved files alone
Name-only evidence files plus an independent checker that re-derives every invariant. A separate advisor model reproduced 2,767 answers from the files alone, and all agreed. Closed PARTIAL: same model family, isolation by instruction only.
1 OctDesign drafts · seven designs, reviewed twice, still DRAFT
Seven designs for the rest of the system. Seven reviews all said "revise". The drafts contradicted each other in 49 places, so one integration contract reconciled them. A follow-up review found most findings addressed and some new ones; the one-follow-up limit stopped the track there (§4).
1 OctThe POC sprint · 8/8 end to end on two clusters
The human pre-approved the recommended options and asked for a working POC. A thin spec cut the designs to the smallest slice that runs end to end; four implementers built it in parallel. 8/8 on Minikube, twice, then 8/8 on a managed cluster created inside the spend cap. An independent review found four important issues; all were reproduced and fixed the same day, and 8/8 passed again on both clusters.
1 OctUI and these pages · a simulator and a read-only live view
A standard-library web server and a single-page simulator on the canonical fixture, with a read-only snapshot of the live services. It runs in the cluster as the seventh core process under a strict content security policy.
1 OctAgents on the data plane · seven task-named agents and a grant-gated store
Six agents train small models on public datasets fetched through grants; a seventh is always denied. The node was upsized one tier (1 vCPU / 2 GiB → 2 vCPU / 4 GiB) under a cap the human raised for it, and 8/8 passed on the new node. A default-deny network policy went in with them.
1 OctPublic UI · the read-only site behind one TLS load balancer
After the human confirmed the DNS plan, the read-only UI was published over TLS only through one connection-throttled cloud load balancer, serving synthetic data. The DNS record was still pending when this page was written.
3What hypotheses changed?
The early working notes warned that the most interesting drift would be architectural, not statistical: an invariant that is correct on its own stops being enough once time, retries, replicas and stale state appear. That is what happened.
| Topic | The notes predicted | Actually decided or built |
|---|---|---|
| Lease size | a weighted demand vector (costs such as 3) | 1 unit per capacity-carrying link; links without a limit cost nothing |
| What admission reserves | runner capacity plus tenant quota | link capacities only; node limits and quotas analysed in-process |
| Partitioning | sampled partitions, epochs, heartbeats, fencing, rebalance | a static disjoint split across 2 partitions, compare both, one fallback on FULL; fencing and rebalance deferred |
| Partition choice | two choices plus a moving average of queue delay | two choices on published spare room, deterministic per request; no moving average |
| Token lifetime | a grant book with renew and revoke | expiry capped 7 s below the lease's; no renew or revoke; a grant may outlive a released lease by ≤ 60 s (disclosed) |
| Internal trust | an enforced mTLS boundary | mTLS deferred; default-deny network policy (managed cluster only); claims scoped to clients that use the front door |
| Exposure | (unspecified) | everything internal except the read-only UI, behind one TLS load balancer |
| Fairness | DRR and batching on the decision path | designed, not built |
| Feedback | PCA → significance → penalty → routing score | PCA maths built; not wired into routing |
| Identity | an external identity provider | its own RS256 issuer; synthetic per-deployment keys; attributes from the catalog |
| Route selection | min-cost flow over a batch | up to 8 legal routes, cheapest first; the first that fits wins |
| Delivery | container images | one digest-pinned official Python image; code in ConfigMaps; no custom image, no registry |
- Ledger safety became distributed ownership. The static split keeps "never overspend" per partition, because no unit has two owners. Review then found the retry hazard the notes predicted (§4).
- Classic DRR became capacity-aware DRR in design review: credit that piles up while nothing fits turns into a future burst, so the draft caps it.
- "Same spectrum" is not a real null. Two matrices with the same eigenvalues can have different eigenvectors, so the drift-penalty design's null test now uses one fixed matrix and two sample seeds.
- Algorithmic versus systems correctness. Max-flow is exact on the canonical graph, while the partition split can still strand capacity (fragmentation).
4What did reviews catch?
Each of these passed its author's own tests and was caught by a reviewer, an oracle or a mutant. Red is before, green after.
An overspend test that passed for the wrong reason
The test failed, but never proved that an overspend would be caught. It was rewritten to fail on the overspend itself.
Jacobi claimed convergence on overflow
The eigenvalue solver returned the unprocessed diagonal labelled "converged". It now says so when it did not converge.
An unbounded search in the evidence checker
A valid but adversarial evidence file could make the checker search for a long time. It now has its own work cap that input cannot raise, and running out is never a PASS.
A loose tolerance gave "OK" with wrong eigenvalues STILL OPEN
The true eigenvalues were 1.874, 0.977 and 0.148. The fix refuses tolerances of 1 or more. A re-check for this page found that 0.5, and 0.49 on another matrix, are still accepted and still return the wrong "OK". Not fixed yet.
A retried request got a second lease
After a lost reply, the retry could be routed to the other partition. The manager now looks the request key up before routing.
A replay after release re-minted a grant
The old OK could be replayed into a fresh grant, or used to release a live lease. An ended lease is never resurrected now.
The judge agreed with NaN
"Is the error larger than the tolerance?" is false for not-a-number, so the comparison helper accepted NaN results.
A distance came back "OK" but NaN
A later threshold check would never have fired on it.
An allow without a scope condition
Without the owner condition, one tenant's rule matched any tenant's agent, which broke tenant isolation.
Slow readers could stall the servers
Every line-JSON server now bounds how long one reply may take to send.
Any pod could skip the front door
Fixed by a default-deny network policy, enforced on the managed cluster (11 of 11 probes as expected). The local cluster ignores it.
Design reviews: better, not perfect
The drafts contradicted each other in 49 places, mapped onto 16 topics in one integration contract, with 8 decisions left for the human. The re-review also flagged 32 places where a draft had settled something that should have been the human's call. The design track then stopped, because one follow-up round is the limit. The POC review of the running code found no critical issue and the four important ones above, all fixed.
5Why were decisions changed?
Most changes of plan came from one of three places: evidence from a test or review, the one-advisor procedure during unattended runs, or a decision by the human.
The disagreement shows why the rule exists. The recommended fix (refuse an empty request rate when loading data) looked clean. The advisor tried it and got 35 test errors, all in policy tests, because policy fixtures legitimately leave the rate empty. The adopted alternative refused the empty value only where a rate is actually used.
Decisions the human made: pre-approving the recommended options for the POC; mutual TLS deferred, so claims are scoped to clients that use the front door; one node size up and a raised spend cap when the agents arrived; fetching the public datasets on the human's own machine, so no dataset credentials ever reach the cluster; a work cap for the evidence checker rather than waiving the check; the PCA tolerance bound; and publishing the read-only UI behind one load balancer.
6What evidence exists?
Hypothesis register
| # | Hypothesis | Status | Evidence |
|---|---|---|---|
| H1 | A stdlib lease ledger admits all-or-nothing across limits and never overspends. | supported, in-process and end to end | 72 tests, 20/20 mutants killed, an oracle check after every call; S3 on both clusters |
| H2 | Graph flow and assignment give exact, checkable limits. | supported on canonical data | independent oracles; the 4 → 2 gap re-derived by an independent checker |
| H3 | A bitmap policy index matches a full scan exactly, and faster. | supported; timing STALE | agreement tests and mutants; 2–21% of scan time on its benchmark, code changed since |
| H4 | Stdlib PCA detects drift with explicit statuses. | maths supported; usefulness untested | oracles, planted spectra, benchmark; tolerance gap open (§4) |
| H5 | Partitions without global per-request writes scale admission. | functional only | 2 partitions, one partition write per admit; no scaling measurement |
| H6 | An independent advisor can verify results from saved logs alone. | supported, PARTIAL | 2,767 advisor answers, all agree; same family, isolation by instruction only |
| H7 | Authenticated, policy-checked, lease-bound path grants work end to end on Kubernetes. | supported (functional POC) | 8/8 on both clusters, also with the workload deployed |
| H8 | Grant-gated data traffic exposes real bottlenecks. | first signal | slow downloads from a CPU-capped store; no capacity contention yet |
Test suites grew with every stage
Tests were written before code, so suite sizes track the work. All bars share one scale.
Data table and where each count comes from
| Member · milestone | Tests | Recorded by |
|---|---|---|
| adaptive-flow · Phase 1 final review | 606 | review record |
| adaptive-flow · after policy bitmaps | 743 | living record |
| adaptive-flow · after PCA maths | 903 | living record |
| adaptive-flow · after log evidence fix wave | 1,057 | living record, re-run at merge |
| adaptive-flow · after PCA config fixes | 1,059 | aggregate verify run |
| adaptive-flow · with policy service | 1,093 | implementer's run (commit message) |
| atomic-lease · baseline / multi-constraint / with servers | 10 / 72 / 138 | living record / verify run / implementer's run |
| authn-sidecar | 86 | commit message |
| POC harness · unit / local end-to-end | 42 / 2 | commit message |
End to end, on two clusters
| Run | Scenarios | Correlated log lines | Trace ids |
|---|---|---|---|
| Minikube, 3 nodes, implementer's run | 8 / 8 PASS | 124 | — |
| Minikube, 3 nodes, coordinator's re-run | 8 / 8 PASS | 226 | 58 |
| Managed cloud cluster, 1 node | 8 / 8 PASS | 123 | 30 |
Every later recorded run, after the review fixes, on the upsized node and with the agent workload deployed, also passed 8 of 8. The clusters use different processor architectures (arm64 locally, amd64 in the cloud) and run the same pinned base image.
The cloud footprint
On the smaller node the seven core pods alone already requested 760m of 940m allocatable CPU (80%), so the node went up one size before the agents were added. Billable extras: one load balancer for the public UI; no volumes.
What the agents measured
A 5-minute window on the managed cluster after warm-up, from the logs of all 15 pods. Times are in milliseconds, median / 95th percentile.
| Agent | Loops | OK | Grant | Verify | Download | Train | MB | Model score (baseline) |
|---|---|---|---|---|---|---|---|---|
agent-credit-default | 21 | 100% | 35.2 / 119 | 4.3 / 17.7 | 3,327 / 5,810 | 4,783 / 9,221 | 60.1 | accuracy 0.8087 (0.7824) |
agent-creditcard-fraud | 22 | 100% | 34.9 / 135 | 3.0 / 11.6 | 3,634 / 7,673 | 4,415 / 5,007 | 232.4 | accuracy 0.9944 (0.9777) |
agent-denied-fleet | 43 | 0% | 104 / 200 | – | – | – | – | 43 of 43 DENIED, no lease, as intended |
agent-german-credit | 36 | 100% | 50.2 / 258 | 3.9 / 21.3 | 68.4 / 266 | 2,203 / 4,804 | 1.8 | accuracy 0.7626 (0.5083) |
agent-mnist | 9 | 100% | 28.4 / 124 | 3.1 / 7.5 | 971 / 2,957 | 23,798 / 28,497 | 12.7 | accuracy 0.8071 (0.1) |
agent-titanic | 40 | 100% | 36.3 / 140 | 4.1 / 9.5 | 70.6 / 104 | 1,398 / 2,605 | 2.4 | accuracy 0.8024 (0.6169) |
agent-wine | 14 | 100% | 39.1 / 129 | 3.1 / 5.8 | 133 / 179 | 13,923 / 19,000 | 1.4 | RMSE 0.6405 (0.8145), lower is better |
The per-agent rows count stats records (142 OK loops); authn and the store count their own log lines (140). The manager logged 140 admits but 142 releases in the window, so two loops had been granted just before it opened. The one extra denial at authn (44 against 43) is likely the same edge effect at the window's end (ASSUMED). Locally, over a 5-minute window of 406 loops, the six allowed agents were 100% OK and the denied agent was refused 67 of 67 times.
- Startup failures before seeding finished. Until the datasets are copied in, the store correctly answers
NOT_SEEDED(19 times on the managed cluster, 8 on Minikube, by its own counters). No bytes are served; the window above starts after warm-up. - The first bottleneck is the download. The 10.6 MB file takes about 3.6 s (median) against 0.8 s locally. The store is held to 100 millicores of CPU and spends about 1.7 s streaming it. More CPU should help, but that has not been tried (ASSUMED).
- Training is about 6–7 times slower than locally (agent-mnist 23.8 s against 3.7 s). Both clusters give each agent the same small CPU limit; the cause has not been isolated (ASSUMED: a slice of a shared cloud vCPU does less work).
- No
FULLcontention yet. Six agents holding at most one lease each never exceed a route with 6 units per link. - Smaller observations. One agent-mnist release found its lease already expired (
NOT_LIVE), because its training comes close to the lease's 30-second maximum; capacity was freed by expiry, not overspent. Signing one assertion takes about 0.7–1.1 s (median) per allowed agent on the managed node, far more than the grant round trip (about 30–50 ms).
Where these numbers come from
The scenario result files and the agent-workload stats files from both clusters (per-agent records, the store's own counters and the manager's overspend check), the living state-of-the-world record, the POC summary, the unattended-run log, the three design-review data files and the specs. Where a count comes only from an implementer's own run (recorded in a commit message), the data table says so. Memory is resident set size, reported in MiB. Figures that could not be checked are marked ASSUMED or left out.
7What did AI agents do?
"Agents" means two different things in this project. The workload agents on the Docs page are small programs that exercise the system. The coding agents below are AI models that built and reviewed it.
| Role | What it did |
|---|---|
| Orchestrator | A strong model planned each stage, split it into tasks, routed tasks to suitable models, and kept the living record up to date. It never switched its own model mid-task. |
| Implementers | Wrote tests first, then code, inside the layering and dependency rules. Four built the POC slice in parallel. |
| Reviewers | Fresh, read-only reviewers for every non-trivial change, one per design, and one for the POC. They found the bugs in §4. |
| Advisors | One fresh advisor per blocker during unattended runs: 9 rounds, 8 AGREE, 1 DISAGREE (§5). A separate advisor reproduced 2,767 evidence answers from saved files alone. |
| Freshness check | A model re-checked 64 claims in the living record and struck none. That is supplementary evidence, not proof. |
| The human | Set the goals, approved consequential choices (§5), held every credential, raised the spend cap once, and is the only one who may push, publish or merge. |
Almost every review was by the same model family as the authors. Agreement there is independence of context, not vendor-independent verification.
8Outcome and what's next
- Authenticated, policy-checked, lease-bound path grants end to end on Kubernetes, on two clusters.
- A continuous workload through the same front door: every allowed loop got data only through an accepted grant, every denied loop held nothing, nothing was overspent.
- In-process: the all-or-nothing ledger, exact flow limits on the canonical data, policy index/scan agreement, the PCA maths.
- Network policy enforced on the managed cluster only; mTLS deferred; the hop check trusts the caller.
- Partitions functional, not shown to scale.
- PCA is maths, not a feature, and its tolerance check is still too loose.
- Fair queuing, drift penalty, rebalance and integration are drafts with 1 critical and 38 important open findings.
- Benchmarks are single-machine; the policy timing is STALE.
- The managed cluster serves the public UI; a teardown has not been recorded.
Next
- Tighten the PCA tolerance bound and add the missing regression test.
- Decide on mutual TLS and real hop attestation now that the network policy is in place.
- Give the dataset store more CPU; add agents or lower capacity until
FULLcontention appears. - Measure partition scaling and fragmentation on the local cluster.
- Get the human's rulings on the pending design decisions; then build fair queuing and wire telemetry into routing.
- Get one cross-family review.
AuthNZ Lab · a proof of concept. All organization names in the fixture are fictional; all keys are synthetic. See Docs for the system itself and the Simulator to explore the canonical fixture.