✳ AuthNZ Lab Policy-constrained execution & partition-owned leases
Journey · how it was built and checked

What was built in three days?

A working path from an agent's request to a check at the destination, running on Kubernetes:

  1. Authenticated agenta signed assertion, replay-proof
  2. Policy graphmay it, and by which legal routes?
  3. Partition-owned capacityall-or-nothing leases, never overspent
  4. Signed path grantnames the lease, the route and its hops
  5. Sidecar verificationat the destination, before any data moves
End-to-end scenarios
8 / 8
pass in every recorded run
Environments
2
local Minikube (3 nodes) and a managed cloud cluster (1 node)
Core processes
7
authn, policy, manager, 2 partitions, sidecar, ui. The workload adds a dataset store and 7 agents: 15 pods.
Core memory
≈ 230 MiB
the 7 core pods together, resident

Here's the unusual part: most of the implementation and review work was done by coding agents operating under a rulebook. A human set the goals, made the consequential calls and held the keys to anything that costs money or leaves the machine.

How the system itself works is on the Docs page. This page is about how it was made, what changed on the way, and what evidence backs it.

1How did we build it?

A file at the root of the repository, the agent working agreement, binds every agent. It is policy, not magic: the project's living record states separately which rules are enforced by code and which are only followed. The agents run inside a harness that routes each job to a suitable model (strong models for hard work and review, light ones for searches and mechanical edits), loads written procedures (skills) when a task needs them, and checks that the living record still matches the code.

Say exactly how sure you are

Every claim carries a label: PASS, FAIL, PARTIAL, BLOCKED, INCONCLUSIVE, STALE or ASSUMED. "It works" is not allowed.

Why: when agents write most of the code, the danger is confident overstatement. A label makes "I didn't check that" visible, so a gap blocks only the claim that depends on it.

Write the failing test first

Watch it fail for the intended reason, then write the smallest code that passes. An import error does not count.

Why: a smoke detector has a test button. A test you have never seen fail may not be connected to anything.

Check every algorithm against an oracle

A second, deliberately simple implementation that shares no code: Dinic against Edmonds–Karp, bitmap policy against a plain scan, Jacobi against power iteration.

Why: to check a sum, add the column a second way. If both ways share a helper, a bug in it is "confirmed" twice.

Name the bug each test should catch

Plant a small deliberate bug (a mutant); the test must fail on it, for the right reason.

Why: a fire drill. Counting alarms proves nothing; lighting a match under one does.

Someone else reviews every change

A fresh, read-only reviewer sees the artifact, the requirements and the evidence, never the author's reasoning. One follow-up round; more need the human.

Why: you read what you meant, not what you wrote. Caveat: reviewers were mostly the same model family, so this is independence of context, not of vendor.

Unattended runs: one second opinion, hard limits stay human

A blocker gets written down with one recommendation and goes to one fresh advisor (AGREE or DISAGREE). Spending beyond a fixed cap, pushing code, DNS, real credentials and weakening a test are never the advisor's call.

Why: a house-sitter can call a plumber but cannot sell the house. One round stops two models talking each other into something.

Two code rules: the runtime uses the Python standard library only, except the member that signs tokens (two pinned, hash-locked cryptography packages that passed a supply-chain check); and imports flow one way through layers, enforced by tests in each member.

2Timeline

One line per phase. Open a line for the detail.

  1. before 29 SepBaseline · a single-constraint lease ledger

    A single-constraint lease ledger with 10 tests and one recorded 14-worker benchmark already existed. The repository history starts from it.

  2. 29–30 SepGraph algorithms · every algorithm against its own oracle

    A 197-row canonical fixture with five fictional customers, then BFS, strongly connected components, topological order, Dijkstra, Dinic max-flow, bounded path enumeration and an exact assignment search. A census records the policy gap (4 → 2). Fitted scaling exponents came out at 0.98–1.00. The whole-branch review covered 40 commits and a 606-test suite: 0 critical, 0 important, 11 should-fix. Of 25 planted mutation probes, 23 were caught; new tests now catch the other 2.

  3. 30 SepPolicy bitmaps · a fast index that must agree with a plain scan

    25 mutants, 0 survivors. Review found one important hole: an allow without the owner-scope condition fired for any tenant's agent. Such rules are now refused at compile time.

  4. 30 SepRules for unattended runs · the night loop, a spend cap, a dependency check

    The rulebook gained its unattended-mode section, the standing cloud authorization with its spend cap, and a security check for any new dependency.

  5. 30 Sep – 1 OctMulti-constraint lease · all-or-nothing across many limits

    Built beside the untouched original so the two can be compared. The lease suite grew from 10 to 72 tests, with 20 of 20 mutants killed.

  6. night → 1 OctFirst unattended run · one controller per track

    Finish the lease, then the PCA maths, then log evidence. Every blocker went through the one-advisor procedure (§5).

  7. 1 OctPCA maths · drift detection as a library

    Windows, scaling, covariance, Jacobi eigenvectors, drift between windows and a whitened distance. Every degenerate case returns a named status. On the long benchmark Jacobi was 19× faster than power iteration at 4 features and 5.2× at 64.

  8. 1 OctLog evidence · results checkable from saved files alone

    Name-only evidence files plus an independent checker that re-derives every invariant. A separate advisor model reproduced 2,767 answers from the files alone, and all agreed. Closed PARTIAL: same model family, isolation by instruction only.

  9. 1 OctDesign drafts · seven designs, reviewed twice, still DRAFT

    Seven designs for the rest of the system. Seven reviews all said "revise". The drafts contradicted each other in 49 places, so one integration contract reconciled them. A follow-up review found most findings addressed and some new ones; the one-follow-up limit stopped the track there (§4).

  10. 1 OctThe POC sprint · 8/8 end to end on two clusters

    The human pre-approved the recommended options and asked for a working POC. A thin spec cut the designs to the smallest slice that runs end to end; four implementers built it in parallel. 8/8 on Minikube, twice, then 8/8 on a managed cluster created inside the spend cap. An independent review found four important issues; all were reproduced and fixed the same day, and 8/8 passed again on both clusters.

  11. 1 OctUI and these pages · a simulator and a read-only live view

    A standard-library web server and a single-page simulator on the canonical fixture, with a read-only snapshot of the live services. It runs in the cluster as the seventh core process under a strict content security policy.

  12. 1 OctAgents on the data plane · seven task-named agents and a grant-gated store

    Six agents train small models on public datasets fetched through grants; a seventh is always denied. The node was upsized one tier (1 vCPU / 2 GiB → 2 vCPU / 4 GiB) under a cap the human raised for it, and 8/8 passed on the new node. A default-deny network policy went in with them.

  13. 1 OctPublic UI · the read-only site behind one TLS load balancer

    After the human confirmed the DNS plan, the read-only UI was published over TLS only through one connection-throttled cloud load balancer, serving synthetic data. The DNS record was still pending when this page was written.

3What hypotheses changed?

The early working notes warned that the most interesting drift would be architectural, not statistical: an invariant that is correct on its own stops being enough once time, retries, replicas and stale state appear. That is what happened.

TopicThe notes predictedActually decided or built
Lease sizea weighted demand vector (costs such as 3)1 unit per capacity-carrying link; links without a limit cost nothing
What admission reservesrunner capacity plus tenant quotalink capacities only; node limits and quotas analysed in-process
Partitioningsampled partitions, epochs, heartbeats, fencing, rebalancea static disjoint split across 2 partitions, compare both, one fallback on FULL; fencing and rebalance deferred
Partition choicetwo choices plus a moving average of queue delaytwo choices on published spare room, deterministic per request; no moving average
Token lifetimea grant book with renew and revokeexpiry capped 7 s below the lease's; no renew or revoke; a grant may outlive a released lease by ≤ 60 s (disclosed)
Internal trustan enforced mTLS boundarymTLS deferred; default-deny network policy (managed cluster only); claims scoped to clients that use the front door
Exposure(unspecified)everything internal except the read-only UI, behind one TLS load balancer
FairnessDRR and batching on the decision pathdesigned, not built
FeedbackPCA → significance → penalty → routing scorePCA maths built; not wired into routing
Identityan external identity providerits own RS256 issuer; synthetic per-deployment keys; attributes from the catalog
Route selectionmin-cost flow over a batchup to 8 legal routes, cheapest first; the first that fits wins
Deliverycontainer imagesone digest-pinned official Python image; code in ConfigMaps; no custom image, no registry
  • Ledger safety became distributed ownership. The static split keeps "never overspend" per partition, because no unit has two owners. Review then found the retry hazard the notes predicted (§4).
  • Classic DRR became capacity-aware DRR in design review: credit that piles up while nothing fits turns into a future burst, so the draft caps it.
  • "Same spectrum" is not a real null. Two matrices with the same eigenvalues can have different eigenvectors, so the drift-penalty design's null test now uses one fixed matrix and two sample seeds.
  • Algorithmic versus systems correctness. Max-flow is exact on the canonical graph, while the partition split can still strand capacity (fragmentation).

4What did reviews catch?

Each of these passed its author's own tests and was caught by a reviewer, an oracle or a mutant. Red is before, green after.

An overspend test that passed for the wrong reason

before mutant: overspend red on another check after mutant: overspend red on the overspend

The test failed, but never proved that an overspend would be caught. It was rewritten to fail on the overspend itself.

caught by adversarial plan review

Jacobi claimed convergence on overflow

before matrix overflows "CONVERGED", raw diag after matrix overflows "UNCONVERGED"

The eigenvalue solver returned the unprocessed diagonal labelled "converged". It now says so when it did not converge.

caught by adversarial review

An unbounded search in the evidence checker

before 36 s on a 356 KB file work grows with input after cap → INCONCLUSIVE input cannot raise it

A valid but adversarial evidence file could make the checker search for a long time. It now has its own work cap that input cannot raise, and running out is never a PASS.

caught by review; cap required by the human

A loose tolerance gave "OK" with wrong eigenvalues STILL OPEN

before tol 0.5, 0 sweeps OK (1, 1, 1) after tol ≥ 1: refused tol 0.49: OK (1, 1, 1)

The true eigenvalues were 1.874, 0.977 and 0.148. The fix refuses tolerances of 1 or more. A re-check for this page found that 0.5, and 0.49 on another matrix, are still accepted and still return the wrong "OK". Not fixed yet.

caught by review; blocker decided by the human

A retried request got a second lease

before k → P0 lease retry k → P1 lease 2 leases after k → P0 lease retry k → found 1 lease

After a lost reply, the retry could be routed to the other partition. The manager now looks the request key up before routing.

caught by the POC review

A replay after release re-minted a grant

before admit k release replay k → OK again after admit k release replay k → ENDED

The old OK could be replayed into a fresh grant, or used to release a live lease. An ended lease is never resurrected now.

caught by the POC review

The judge agreed with NaN

before NaN > tol → false counted as a match after not (NaN ≤ tol) counted as a failure

"Is the error larger than the tolerance?" is false for not-a-number, so the comparison helper accepted NaN results.

caught by task review

A distance came back "OK" but NaN

before 1e10 → ∞ → NaN status OK, value NaN after 1e10 → not finite error names feature

A later threshold check would never have fired on it.

caught by review

An allow without a scope condition

before Cedar allow, no scope fires for Northstar after Cedar allow, no scope refused at compile

Without the owner condition, one tenant's rule matched any tenant's agent, which broke tenant isolation.

caught by the final review

Slow readers could stall the servers

before client stops reading server writer stuck after client stops reading write timeout, closed

Every line-JSON server now bounds how long one reply may take to send.

caught by the POC review

Any pod could skip the front door

before any pod → manager, skips authn after any pod blocked (managed cluster)

Fixed by a default-deny network policy, enforced on the managed cluster (11 of 11 probes as expected). The local cluster ignores it.

caught by the POC review

Design reviews: better, not perfect

First review · 7 designs
114
findings: 21 critical, 64 important, 29 minor, plus 21 wrong citations. 7 of 7 said "revise".
Triage
137 / 1 / 39
review items fixed / rejected with evidence / left unresolved as decisions
Re-review
1 / 38 / 45
new critical / important / minor. Earlier findings: 125 addressed, 22 partly, 0 ignored, 0 wrongly rejected.

The drafts contradicted each other in 49 places, mapped onto 16 topics in one integration contract, with 8 decisions left for the human. The re-review also flagged 32 places where a draft had settled something that should have been the human's call. The design track then stopped, because one follow-up round is the limit. The POC review of the running code found no critical issue and the four important ones above, all fixed.

5Why were decisions changed?

Most changes of plan came from one of three places: evidence from a test or review, the one-advisor procedure during unattended runs, or a decision by the human.

Advisory rounds
9
3 for the PCA maths, 5 for log evidence, 1 for the design conflicts
AGREE
8
facts re-checked against source before acting, then verified by tests
DISAGREE
1
the advisor's alternative was right and was adopted

The disagreement shows why the rule exists. The recommended fix (refuse an empty request rate when loading data) looked clean. The advisor tried it and got 35 test errors, all in policy tests, because policy fixtures legitimately leave the rate empty. The adopted alternative refused the empty value only where a rate is actually used.

Decisions the human made: pre-approving the recommended options for the POC; mutual TLS deferred, so claims are scoped to clients that use the front door; one node size up and a raised spend cap when the agents arrived; fetching the public datasets on the human's own machine, so no dataset credentials ever reach the cluster; a work cap for the evidence checker rather than waiving the check; the PCA tolerance bound; and publishing the read-only UI behind one load balancer.

6What evidence exists?

Hypothesis register

#HypothesisStatusEvidence
H1A stdlib lease ledger admits all-or-nothing across limits and never overspends.supported, in-process and end to end72 tests, 20/20 mutants killed, an oracle check after every call; S3 on both clusters
H2Graph flow and assignment give exact, checkable limits.supported on canonical dataindependent oracles; the 4 → 2 gap re-derived by an independent checker
H3A bitmap policy index matches a full scan exactly, and faster.supported; timing STALEagreement tests and mutants; 2–21% of scan time on its benchmark, code changed since
H4Stdlib PCA detects drift with explicit statuses.maths supported; usefulness untestedoracles, planted spectra, benchmark; tolerance gap open (§4)
H5Partitions without global per-request writes scale admission.functional only2 partitions, one partition write per admit; no scaling measurement
H6An independent advisor can verify results from saved logs alone.supported, PARTIAL2,767 advisor answers, all agree; same family, isolation by instruction only
H7Authenticated, policy-checked, lease-bound path grants work end to end on Kubernetes.supported (functional POC)8/8 on both clusters, also with the workload deployed
H8Grant-gated data traffic exposes real bottlenecks.first signalslow downloads from a CPU-capped store; no capacity contention yet

Test suites grew with every stage

Tests were written before code, so suite sizes track the work. All bars share one scale.

Test suite size at each milestone, per code member adaptive-flow · Phase 1 review 606 tests606 adaptive-flow · after policy bitmaps 743 tests743 adaptive-flow · after PCA maths 903 tests903 adaptive-flow · after log evidence 1,057 tests1,057 adaptive-flow · after PCA config fixes 1,059 tests1,059 adaptive-flow · with policy service 1,093 tests1,093 atomic-lease · baseline 10 tests10 atomic-lease · multi-constraint 72 tests72 atomic-lease · with servers 138 tests138 authn-sidecar · new member 86 tests86
Data table and where each count comes from
Member · milestoneTestsRecorded by
adaptive-flow · Phase 1 final review606review record
adaptive-flow · after policy bitmaps743living record
adaptive-flow · after PCA maths903living record
adaptive-flow · after log evidence fix wave1,057living record, re-run at merge
adaptive-flow · after PCA config fixes1,059aggregate verify run
adaptive-flow · with policy service1,093implementer's run (commit message)
atomic-lease · baseline / multi-constraint / with servers10 / 72 / 138living record / verify run / implementer's run
authn-sidecar86commit message
POC harness · unit / local end-to-end42 / 2commit message

End to end, on two clusters

RunScenariosCorrelated log linesTrace ids
Minikube, 3 nodes, implementer's run8 / 8 PASS124—
Minikube, 3 nodes, coordinator's re-run8 / 8 PASS22658
Managed cloud cluster, 1 node8 / 8 PASS12330

Every later recorded run, after the review fixes, on the upsized node and with the agent workload deployed, also passed 8 of 8. The clusters use different processor architectures (arm64 locally, amd64 in the cloud) and run the same pinned base image.

The cloud footprint

Core stack
≈ 230 MiB
the 7 core pods together (235,904 KiB resident), on the smaller node, before the agents
With the workload
860m · 902 MiB
CPU and memory requested by all 15 pods: 44% and 32% of the 2 vCPU / 4 GiB node
Per pod in use
27–50 MiB
sampled resident memory, for example dataset-store ≈ 27, authn ≈ 45

On the smaller node the seven core pods alone already requested 760m of 940m allocatable CPU (80%), so the node went up one size before the agents were added. Billable extras: one load balancer for the public UI; no volumes.

What the agents measured

A 5-minute window on the managed cluster after warm-up, from the logs of all 15 pods. Times are in milliseconds, median / 95th percentile.

AgentLoopsOKGrantVerifyDownloadTrainMBModel score (baseline)
agent-credit-default21100%35.2 / 1194.3 / 17.73,327 / 5,8104,783 / 9,22160.1accuracy 0.8087 (0.7824)
agent-creditcard-fraud22100%34.9 / 1353.0 / 11.63,634 / 7,6734,415 / 5,007232.4accuracy 0.9944 (0.9777)
agent-denied-fleet430%104 / 200––––43 of 43 DENIED, no lease, as intended
agent-german-credit36100%50.2 / 2583.9 / 21.368.4 / 2662,203 / 4,8041.8accuracy 0.7626 (0.5083)
agent-mnist9100%28.4 / 1243.1 / 7.5971 / 2,95723,798 / 28,49712.7accuracy 0.8071 (0.1)
agent-titanic40100%36.3 / 1404.1 / 9.570.6 / 1041,398 / 2,6052.4accuracy 0.8024 (0.6169)
agent-wine14100%39.1 / 1293.1 / 5.8133 / 17913,923 / 19,0001.4RMSE 0.6405 (0.8145), lower is better
Front door
140 OK · 44 denied
grants issued and refused by authn in the window
Dataset store
140 served
every download had a grant the sidecar accepted
Overspend
none
checked against the manager's snapshot of every limit

The per-agent rows count stats records (142 OK loops); authn and the store count their own log lines (140). The manager logged 140 admits but 142 releases in the window, so two loops had been granted just before it opened. The one extra denial at authn (44 against 43) is likely the same edge effect at the window's end (ASSUMED). Locally, over a 5-minute window of 406 loops, the six allowed agents were 100% OK and the denied agent was refused 67 of 67 times.

  • Startup failures before seeding finished. Until the datasets are copied in, the store correctly answers NOT_SEEDED (19 times on the managed cluster, 8 on Minikube, by its own counters). No bytes are served; the window above starts after warm-up.
  • The first bottleneck is the download. The 10.6 MB file takes about 3.6 s (median) against 0.8 s locally. The store is held to 100 millicores of CPU and spends about 1.7 s streaming it. More CPU should help, but that has not been tried (ASSUMED).
  • Training is about 6–7 times slower than locally (agent-mnist 23.8 s against 3.7 s). Both clusters give each agent the same small CPU limit; the cause has not been isolated (ASSUMED: a slice of a shared cloud vCPU does less work).
  • No FULL contention yet. Six agents holding at most one lease each never exceed a route with 6 units per link.
  • Smaller observations. One agent-mnist release found its lease already expired (NOT_LIVE), because its training comes close to the lease's 30-second maximum; capacity was freed by expiry, not overspent. Signing one assertion takes about 0.7–1.1 s (median) per allowed agent on the managed node, far more than the grant round trip (about 30–50 ms).

Where these numbers come from

The scenario result files and the agent-workload stats files from both clusters (per-agent records, the store's own counters and the manager's overspend check), the living state-of-the-world record, the POC summary, the unattended-run log, the three design-review data files and the specs. Where a count comes only from an implementer's own run (recorded in a commit message), the data table says so. Memory is resident set size, reported in MiB. Figures that could not be checked are marked ASSUMED or left out.

7What did AI agents do?

"Agents" means two different things in this project. The workload agents on the Docs page are small programs that exercise the system. The coding agents below are AI models that built and reviewed it.

RoleWhat it did
OrchestratorA strong model planned each stage, split it into tasks, routed tasks to suitable models, and kept the living record up to date. It never switched its own model mid-task.
ImplementersWrote tests first, then code, inside the layering and dependency rules. Four built the POC slice in parallel.
ReviewersFresh, read-only reviewers for every non-trivial change, one per design, and one for the POC. They found the bugs in §4.
AdvisorsOne fresh advisor per blocker during unattended runs: 9 rounds, 8 AGREE, 1 DISAGREE (§5). A separate advisor reproduced 2,767 evidence answers from saved files alone.
Freshness checkA model re-checked 64 claims in the living record and struck none. That is supplementary evidence, not proof.
The humanSet the goals, approved consequential choices (§5), held every credential, raised the spend cap once, and is the only one who may push, publish or merge.

Almost every review was by the same model family as the authors. Agreement there is independence of context, not vendor-independent verification.

8Outcome and what's next

Supported
  • Authenticated, policy-checked, lease-bound path grants end to end on Kubernetes, on two clusters.
  • A continuous workload through the same front door: every allowed loop got data only through an accepted grant, every denied loop held nothing, nothing was overspent.
  • In-process: the all-or-nothing ledger, exact flow limits on the canonical data, policy index/scan agreement, the PCA maths.
Partial or deferred
  • Network policy enforced on the managed cluster only; mTLS deferred; the hop check trusts the caller.
  • Partitions functional, not shown to scale.
  • PCA is maths, not a feature, and its tolerance check is still too loose.
  • Fair queuing, drift penalty, rebalance and integration are drafts with 1 critical and 38 important open findings.
  • Benchmarks are single-machine; the policy timing is STALE.
  • The managed cluster serves the public UI; a teardown has not been recorded.

Next

  1. Tighten the PCA tolerance bound and add the missing regression test.
  2. Decide on mutual TLS and real hop attestation now that the network policy is in place.
  3. Give the dataset store more CPU; add agents or lower capacity until FULL contention appears.
  4. Measure partition scaling and fragmentation on the local cluster.
  5. Get the human's rulings on the pending design decisions; then build fair queuing and wire telemetry into routing.
  6. Get one cross-family review.

AuthNZ Lab · a proof of concept. All organization names in the fixture are fictional; all keys are synthetic. See Docs for the system itself and the Simulator to explore the canonical fixture.