Inside FoundryOS

Summary

During one of my Cure runs, an investigator reported three findings, including one it described as P0.

The persisted artifact contained no findings.

The statement and the evidence disagreed. The operation refused to advance. The missing artifact was not treated as a formatting inconvenience: it meant the downstream aggregator had nothing trustworthy to consume. The work was redispatched and re-evidenced before anything moved.

That refusal is the character of the system this document describes.

FoundryOS does not ask one model to pretend it is an entire company.

FoundryOS is an operator-governed, automated software engineering department that turns business intent into built, provisioned, verified, released and continuously improving software. It is not one model, one agent, one prompt or one linear pipeline. It is a set of bounded responsibilities joined by persistent state, explicit authority, deterministic controls, separate verification paths and a learning loop that changes the system after preventable failure.

This is the technical companion to FoundryOS: An Automated Software Engineering Department. The foundational article explains the category and the business problem. This reference explains the mechanisms, the evidence, the boundaries and the operating history. It was not designed all at once on a whiteboard. Production work forged it, and this document shows the forge marks where they happened.

30-SECOND ARCHITECTURE

The nine-step end-to-end flow of the department: business intent, operational truth, canonical build authority, persistent project context, governed execution, construction and infrastructure provisioning, plural verification, controlled release, maintenance and institutional learning. Models provide intelligence; the department supplies ownership, state, verification and consequence control.

FoundryOS is organized like an automated software engineering department.

Models provide intelligence. The department around that intelligence supplies clear ownership, persistent state, bounded authority, isolated execution, infrastructure automation, separate verification, consequence-aware release controls and durable project and fleet knowledge.

The central engineering rule is:

Principle: Fix the failure. Then fix the system that allowed it.

HOW TO READ THIS

This is a living architecture reference, not a claim that every task invokes every component.

For a business reader: Read the 30-second architecture, the system map, the early CrewOS trace, Sections 11 and 12, and the conclusion.

For an engineering reader: Follow the categories in order. Each begins with a contract, identifies its mechanisms and explains why the boundary exists. Where a recorded failure produced the boundary, the failure leads the section.

For a governance or audit reader: Start with Evidence Posture, the Failure-Routing Contract, Verification and Remediation, Release and Consequence Control, the Operational Failure and Control Model, and Open Engineering Problems.

For a skeptical reader: Section 14 contains the five questions for attacking any claim in this document: contract, mechanism, fixture, run artifact and evidence class. You are invited to start there and work backward.

For a FoundryOS operator: Use the contents and the component reference as navigation. Contracts, ownership boundaries and evidence returns are written to support operating use as well as public explanation.

One build runs through the middle of this document. Each of the seven responsibility categories ends with what that responsibility looked like in the CrewOS external-participant build, and Section 11 tells the build whole.

1. EVIDENCE POSTURE

A seven-rung evidence ladder: 1 installed, 2 mechanically enforced (hooks, tests, schemas or state machines), 3 exercised in operating work, 4 outcome evidenced (a documented catch, recovery, correction or release), 5 fleet-codified (a local lesson became a shared standard or reusable control), 6 comparative evidence, 7 portability evidence (exercised across a different operator, stack, model, provider or environment).

Contract: Read every FoundryOS claim according to the evidence that supports it, not according to how ambitious the surrounding language sounds.

The useful distinction is not whether a mechanism exists. It is what kind of evidence supports the claim being made about it.

Evidence level What it establishes
Installed The mechanism exists in the current system
Mechanically enforced Tests, hooks, schemas, state machines, scanners or criteria protect its core contract
Exercised in operating work The mechanism has participated in real FoundryOS or application work
Outcome evidenced A documented catch, recovery, correction, certification or release exists
Fleet-codified A local lesson became a shared standard, reusable pattern or keep-installed control
Comparative evidence The mechanism or outcome has been compared with another approach under controlled or meaningfully comparable conditions
Portability evidence The mechanism has been exercised across a different operator, technology stack, model, provider or transport environment

FoundryOS has strong installed, mechanically enforced and operating evidence across much of the architecture. It also has outcome evidence for specific incidents and fleet codification for many lessons.

The weaker areas are comparative economics, cross-operator portability, long-term recurrence measurement and broad cross-stack replication. Comparative performance and portability are separate claims, and neither should be inferred from the other. Those remain engineering and evidence problems.

Current implementation snapshot

The values below were rechecked against current source on August 2, 2026, against worktree HEAD 24ec65306. They must be regenerated from source on the publication date.

Crucible candidates pass through controlled collection, filtering, categorization and deduplication before distribution. Source context is retained and unresolved contradictions remain explicit. Converge analyzes catches and proposes codification; control retirement remains an explicit operator-and-plan action. The count therefore measures enumerated knowledge entries, not equal evidentiary weight or universal applicability.

The Caliber count measures registered shipped criteria contributing to the current scale. Criteria differ in severity, blocking behaviour, fixture depth and operating history. The evidence package should publish separate denominators for criteria with red-and-green fixtures, criteria that caught an operating defect, advisory versus blocking criteria, retired criteria and reviewed false positives.

These are implementation and operating facts. They do not establish a universal productivity multiplier or a complete replacement for every human engineering function.

Evidence by selected mechanism

The classifications below are deliberately conservative.

Mechanism Installed Enforced Exercised Outcome evidence Fleet-codified Comparative evidence Portability evidence
DO NOT STOP Yes Yes Yes Yes Yes Not demonstrated Partial
Chassis Yes Yes Yes Yes Not applicable Not demonstrated Partial
Auto Plan Yes Yes Yes Yes Partial Not demonstrated Partial
Cure Yes Yes Yes Yes Yes Not demonstrated Partial
Caliber Yes Yes Yes Yes Yes Not demonstrated Partial
Codex audit Yes Yes Yes Yes Partial Not demonstrated Partial
Caster and Crucible Yes Yes Yes Yes Yes Not demonstrated Partial
Converge Yes Partial Yes Early Partial Not demonstrated Not demonstrated
Root-operator delegation Partial Partial Limited Limited No Not demonstrated Not demonstrated

The most important column in that table is the one that reads Not demonstrated from top to bottom. It is there on purpose. A table that admitted no gaps would tell you nothing about the entries that claim strength.

In this table, comparative evidence means a controlled or meaningfully comparable evaluation against another approach. Portability evidence means execution across a different operator, stack, model, provider or transport environment.

A reader can challenge any claim by asking for the contract, implementation, fixture, operating artifact and evidence class. Section 14 formalizes that challenge path.

2. COMPRESSED SYSTEM MAP

System intent flows through seven responsibility categories to delivered value: business definition and authority (Charge, Channel); foundation and persistent context (Core, Codex Template, Project Codex, Crucible); execution and continuity (DO NOT STOP, go codex audit, Chassis profiles); build and provisioning (Conductor, Constructor, Converter, Cell, Commission); verification and remediation (Cure, Caliber, Codex audit, Completer, Commission verification); release and consequence control (Close, Promote, Prune, Deploy); maintenance and learning (Caster, Crucible, Converge, Caliber).

Contract: Every part of the department must have a bounded responsibility, an authority limit and an evidence return.

FoundryOS is easiest to understand as seven responsibility categories.

Responsibility category Purpose Primary mechanisms
Business Definition and Authority Convert operational reality into authorized work Charge, Channel, business authority, root operator
Foundation and Persistent Context Give workers architecture, memory and reusable doctrine Core, Codex Template, Project Codex, Crucible
Execution and Continuity Carry work through legitimate progression without false completion DO NOT STOP, go codex audit, Chassis, Auto Plan, Cure profiles
Build and Provisioning Construct the capability and connect it to external infrastructure Conductor, Constructor, Converter, Cell, Commission, worktrees
Verification and Remediation Challenge the result from distinct perspectives and repair defects Cure, Caliber, Codex audit, Completer, Commission verification
Release and Consequence Control Separate completion, preservation, integration and deployment Close, promote, prune, deploy, rollback and destructive-change gates
Maintenance and Institutional Learning Keep knowledge current and convert catches into durable controls Caster, Crucible, Converge, Project Codex, Caliber criteria

The 15 C-components occupy those categories, but the architecture also includes shared machinery that is not itself a C-component. Chassis, Auto Plan, DO NOT STOP, the Codex audit path, hooks, schemas, ledgers and release scripts are examples.

Compressed component map

Component Departmental role
Charge Intake and operational discovery
Channel Canonical specification and decision authority
Core Standard application foundation and shared module catalogue
Codex Template Standard project-memory structure
Crucible Reusable fleet knowledge
Conductor Lifecycle sequencing, bounded decomposition, coordination and integration readiness
Constructor Governed application construction inside approved scope
Completer Promise-to-implementation completeness
Converter Existing-system conversion and reusable migration doctrine
Cell Mobile, PWA and store readiness
Commission Infrastructure and provider provisioning
Cure Running-system investigation, repair and certification
Caliber Codified structural quality governance
Caster Project and fleet knowledge maintenance and distribution
Converge Operator-catch analysis, codification proposals and control-gap surfacing

One real build in one minute

The complete CrewOS account appears in Section 11. Its compressed operating trace is:

The trace matters because it shows the architecture moving real work, including a deliberate halt when source identity could not support a trustworthy verdict.

3. FAILURE-ROUTING CONTRACT

The failure-routing contract in seven steps: failure, observation point, detection owner, repair or decision owner, required evidence, permitted next state, and human authority where necessary. A failure without an owner becomes operator memory; a failure with evidence but no return route becomes a dead end.

Contract: Every meaningful failure needs an observation point, an owner, a remedy, required evidence and a permitted next state.

Start with one failure walked all the way through, because the grammar is easier to trust after you have seen it hold real weight.

Walked example: unsafe cleanup after partial promotion

During Cure v2 blind-benchmark work, a promotion path reached PARTIAL_SUCCESS and still entered cleanup. An orphaned submodule commit was destroyed before it could be pushed somewhere reachable. I lost work. The full release mechanism appears in Section 9; here is the incident expressed through the routing grammar:

Routing field Bound value
Failure Cleanup ran after partial promotion and destroyed work that had not been safely preserved
Observation point The prune boundary
Detection owner Prune preflight and the C112 keep-installed control
Repair owner Release-control substrate
Required evidence Source reachability and preservation proof before deletion
Permitted next state Proceed only after preservation is proven; halt on ambiguity
Human authority Required when source identity or preservation cannot be established mechanically

This is the difference between documenting a bad outcome and binding the system against its recurrence.

The grammar

The routing model is the binding grammar of the system. If a failure class is absent from it, the failure remains an operator catch by default. Every route answers three practical questions: Who detects it? Who owns the response? What must be true before progression resumes?

The other thirteen

Failure class First detector Repair or decision route Progression requirement
Missing business requirement Charge or Channel Return to business authority Canonical decision exists
Conflicting source authority Channel Must-resolve gate Source precedence is explicit
Invalid plan premise Design review or Codex audit Revise or elevate Objective and architecture remain valid
Premature model stop DO NOT STOP or Chassis Continue or authorized elevation Legitimate terminal state exists
Out-of-scope file mutation Path guard or integration review Halt and isolate Authorized path ownership is restored
Parallel implementation collision Worktree and integration controls Reconcile through integrator Semantic conflict is resolved
Migration written but unapplied Migration hook, Commission or Cure Apply and verify Live schema matches required state
Missing provider configuration Commission preflight Provision or structured human gate External state is verified
Runtime defect Cure Investigate, repair and retest Running behaviour passes again
Known structural violation Caliber Fix or authorized waiver Criterion no longer blocks
Novel design or security issue Codex audit Fold, fix and re-audit Finding is resolved or elevated
Missing promised capability Completer Completion plan Specification and implementation agree
Unsafe cleanup Prune gate Preserve and establish reachability No unpublished work can be lost
Repeated operator rescue Converge Codification proposal or operator-routed control review Prevention is implemented, declined or routed for explicit operator action

A failure without an owner becomes operator memory. A failure with an owner but no evidence contract becomes an opinion. A failure with evidence but no return route becomes a dead end.

4. BUSINESS DEFINITION AND AUTHORITY

Capability is not authority: the governing chain runs capability, component contract, explicit permission, returned evidence, authorized consequence. A model can identify ambiguity in policy but cannot authorize what the policy should become.

Contract: Engineering begins only after business reality has been converted into explicit build authority.

Why this boundary exists. Early implementation can be technically competent and still encode the wrong policy. Business authority must be settled before builders convert assumptions into software.

Business authority

The owner, operator, subject-matter expert or client stakeholder owns:

A model can identify ambiguity in a cancellation policy. It cannot authorize what the policy should become.

Charge

Charge owns intake and operational discovery.

It begins with the material businesses actually have:

Its job is to discover the operation behind the requested software.

Charge identifies actors, records, handoffs, exceptions, constraints, sources of truth, recurring failures and unresolved decisions. It does not build the application.

The output is structured business context that another component can challenge and formalize.

Channel

Channel converts business context into canonical specification.

It reconciles contradictory sources, records decisions, defines roles and permissions, maps workflows, establishes acceptance conditions and blocks on unresolved business authority.

The important output is not a long requirements document. It is build authority.

Constructor should not invent policy. Commission should not infer provider ownership. Cure should not decide that an incomplete promised workflow is acceptable. Those decisions have to exist before the downstream component acts.

Root operator

I currently hold the root-operator role.

The root operator routes work across the department, holds provider-account authority, resolves consequential gates, manages integration and remains accountable for the system. The architecture separates that authority from implementation so delegation can become explicit and testable. Cross-operator replication has not yet been demonstrated.

The operator should not quietly absorb every other role. That would recreate the single-context failure FoundryOS is designed to avoid.

Capability is not authority

The governing chain is:

A fixer may be capable of deploying to production and forbidden to do so. A reviewer may understand the repair and be prohibited from modifying the target it grades. An infrastructure component may prepare a billing integration and still require human authorization for a financial commitment.

In the CrewOS build: the external-participant capability was represented as a plan, not a collection of prompts. Scope, ownership, acceptance conditions and touched surfaces were explicit before broad implementation began.

5. FOUNDATION AND PERSISTENT CONTEXT

Contract: A worker enters an established engineering environment, not an empty repository plus a large prompt.

Why this boundary exists. Session memory is temporary, summaries can drift and workers can inherit incomplete assumptions. Persistent context and a truth hierarchy prevent every session from reconstructing the project from scratch.

Core

Core is the standard application foundation, framework doctrine, reusable module catalogue and owner of shared Chassis machinery.

It supplies repeatable architecture and implementation patterns so each application does not have to rediscover:

Core is not a single generated starter. It is the maintained technical foundation from which compatible applications and FoundryOS mechanisms can begin.

Codex Template and Project Codex

The Codex Template defines the standard project-memory structure.

A Project Codex records the local truth another worker needs:

The Project Codex is not allowed to overrule the live system. It is maintained context subordinate to observed reality and canonical source.

Crucible

Crucible contains reusable knowledge that should survive beyond one application.

It holds generalized lessons, patterns, warnings and implementation doctrine that future projects can consume when the stack and situation match.

Other components can identify Crucible candidates. Caster owns controlled writes and distribution so shared knowledge is not modified ad hoc by every worker.

Truth hierarchy

FoundryOS uses a conceptual truth hierarchy:

  1. Running system and observed external reality.
  2. Binding framework rules and canonical source.
  3. Crucible reusable knowledge.
  4. Project documentation and Project Codex.
  5. Generated projections and indexes.

A model summary cannot overrule the artifact it claims to summarize. A Project Codex statement cannot overrule live schema inspection. A generated count cannot overrule mechanically enumerated source.

Caster Spin Up

A project is initialized from the combination of:

Core
+ Codex Template
+ relevant Crucible knowledge
+ project-specific business authority

This gives a new worker a technical foundation, a memory structure and stack-matched lessons before substantial implementation begins.

In the CrewOS build: workers entered a project with a maintained Project Codex, and Caster refreshed that knowledge after the build so the next worker inherits the current truth rather than a reconstruction.

6. EXECUTION AND CONTINUITY

Chassis is a reusable execution framework (persistent run state, stage progression, retries and budgets, decision routing, handoff recovery, terminal states) with three consumer profiles: Auto Plan (complex planned engineering work), Cure v2 (structured investigation, remediation and certification), and Cure v3 fix-pass (specialized automated repair). A profile supplies declarative data, seam wiring and procedure. Chassis supplies motion, state and control; the profile supplies the route.

Four FoundryOS routes matched to work shape: a small bounded list uses DO NOT STOP; small or ad hoc work needing review uses DO NOT STOP plus go codex audit; a complex multi-stage engineering plan uses Auto Plan on Chassis; structured running-system investigation and repair uses Cure v2 or v3 on Chassis.

Contract: FoundryOS applies enough execution structure to preserve completion, evidence and consequence control without forcing minor work through maximum ceremony.

Why this boundary exists. In one of my production sessions, Claude wrote that 225 cells remained and said it was continuing. Then it stopped. The instruction was already clear. Nothing enforced it at the boundary where the turn ended. A thin prompt can lose a consequential program inside one context window, while a heavyweight lifecycle can drown a five-line correction in process. Execution weight must match the shape and consequence of the work.

This category contains the most frequently confused parts of the architecture.

Chassis

Chassis is FoundryOS's reusable, profile-driven execution framework for carrying consequential work through a durable lifecycle.

Principle: Chassis supplies motion, state and control. The profile supplies the route.

The generic engine owns machinery such as:

The engine does not define the meaning of a particular lifecycle.

A consumer profile supplies:

  1. Declarative data: stages, exit enums, forks, routing rows, consequence maps, budgets and model contracts.
  2. Seam wiring: consumer-specific implementations plugged into the generic engine.
  3. Procedure: the handlers that execute the consumer's actual stages.

The current engine has three active consumers:

This distinction matters. Chassis is not Auto Plan. Auto Plan is one route expressed on Chassis.

Auto Plan

Auto Plan is the lifecycle for significant engineering work that needs durable planning, review, implementation, correction and closure across component boundaries.

Its current simplified profile is expressed as eight lifecycle stages:

The profile carries a plan from initial authority through separate Codex review, construction, finding resolution and a legitimate terminal state.

Components still own their bounded responsibilities. Charge owns discovery. Constructor owns application construction. Commission owns infrastructure. Cure owns running-system investigation. Caliber owns codified structural grading.

Auto Plan defines how a complex engineering plan progresses across those responsibilities. Chassis supplies the reusable execution framework underneath it.

Cure v2 and Cure v3

Cure uses Chassis for a different kind of work.

The Cure profile carries a running-system investigation through scope, investigation, aggregation, repair, retest, lifecycle reconciliation and certification or elevation.

The Cure v2 operation is represented as a validated gate graph with persistent state rather than as a checklist the operator must remember. Its detailed structural counts are retained in the implementation notes rather than treated as outcome metrics.

Cure v3 uses the same generic Chassis foundation for specialized fix-pass execution.

DO NOT STOP

DO NOT STOP is the lightweight execution contract. It exists because of sessions like the one that opened this section.

It covers bounded work where the main risks are premature stopping, dropped list items and false completion.

The mechanism binds the lifecycle event where stopping becomes observable. Repeating "do not stop" in a prompt is not enough if the runtime still attempts to terminate the turn.

DO NOT STOP does not provide the full Auto Plan lifecycle. It provides continuation and completion discipline for smaller work.

go codex audit

go codex audit is the canonical separate Codex-review route for work not driven through Auto Plan.

It supports:

The driver resolves scope, performs readiness checks, runs the correct audit route, fails loudly with a reason and remedy, and is enforced as the sanctioned non-Chassis entry path.

It does not replace Auto Plan and it does not replace the post-land deferred-audit safety net.

Execution weight

The architecture is proportional by design.

Each layer protects a distinct invariant:

Two-signal surfacing

An active Chassis run should not casually fall out of its lifecycle and ask the operator an off-contract question.

The two-signal contract binds the session to a non-terminal run and blocks ordinary surfacing until the run reaches a real terminal state or a logged elevation legitimately returns authority.

This converts "keep going" from etiquette into an execution property.

In the CrewOS build: persistent Chassis state carried the planned work across individual agent contexts while retaining stage, retry, decision and handoff information.

7. BUILD AND PROVISIONING

Contract: Approved intent becomes an integrated operational capability, not a pile of generated files.

Why this boundary exists. Parallel builders can produce locally correct parts that disagree on shared types, migrations, ownership and external state. Construction, integration and provisioning therefore remain distinct responsibilities.

Conductor

Conductor sequences work across lifecycle phases, decomposes significant work into bounded packages, coordinates parallelism and owns integration readiness.

Parallel agents can increase throughput. They also create risks:

Conductor's job is not merely to create more workers. It is to keep their work reconcilable.

Constructor

Constructor owns governed application construction inside the approved build phase. It does not define business policy or the cross-lifecycle route.

It receives canonical specification, project context, an approved plan, permitted paths, expected artifacts and verification requirements.

Constructor can use specialized builders and isolated worktrees. The builder is allowed to reason, design and implement inside the contract. It is not allowed to redefine the business or silently expand its authority.

Worktrees and isolation

FoundryOS uses isolated worktrees for parallel implementation where appropriate.

Isolation creates a separate working surface for each bounded task. It does not solve integration automatically.

A worktree can be clean and still contain an implementation based on the wrong premise. Two worktrees can both pass tests and still disagree semantically. A branch can contain valuable work that is not reachable from a protected remote branch.

FoundryOS therefore treats integration and source reachability as distinct responsibilities.

Principle: Parallelism produces parts. Integration produces a system.

Converter

Converter turns an existing application or operational system into a FoundryOS-compatible product without discarding the business knowledge already embedded in it.

The scar here is recorded: after the Valley Mainland conversion, sub-users could not resolve the account they belonged to. A production system carries identity assumptions that a greenfield scaffold never meets. The Valley Mainland to CrewOS conversion produced reusable doctrine around:

Converter exists because a production application carries assumptions that a greenfield scaffold does not.

Cell

Cell owns mobile, PWA and app-store readiness.

Its scar is also recorded: mobile login tests once stopped before reaching authenticated behaviour, passing on a signal that proved nothing. A responsive web interface is not automatically mobile-ready. Cell addresses touch behaviour, viewport constraints, authenticated mobile paths, accessibility, performance, PWA configuration, native wrappers, store assets, privacy disclosures, signing and distribution requirements.

Cell is invoked where the project's registered targets require it. It is not a mandatory stage for every application.

Commission

Commission owns infrastructure and provider provisioning.

Its surface can include:

Commission distinguishes source truth from external truth.

A migration file in Git does not prove the migration reached production. A webhook configuration in code does not prove the provider registered it. A deployment manifest does not prove DNS points to the deployed system.

The infrastructure component has to inspect or exercise the external system.

Consequence-aware construction

Not every implementation action has the same blast radius.

A local edit, committed migration, live schema change, outbound customer message, billing action and production deployment carry different consequences.

FoundryOS uses permissions, gates and escalation according to reversibility and authority rather than model confidence alone.

In the CrewOS build: the operating record reports 146 commits progressing across multiple bounded worktrees, with integration held as a separate responsibility because clean branches can still disagree on authorization semantics, shared types and source state.

8. VERIFICATION AND REMEDIATION

Contract: The implementing agent contributes evidence, but it does not control the final verdict.

Why this boundary exists. Builders naturally inspect the work through the assumptions that produced it, and verification machinery can fail in its own right: a repository guard once returned PASS while inspecting the wrong root. Separate verification domains are needed because runtime, structure, promised scope and external state can fail independently.

Correctness is plural.

Verification system Primary question
Cure Does the running system behave correctly across the required surfaces?
Caliber Does the project retain known structural and process controls?
Codex audit Can a role-, context- and source-separated reviewer find a defect or flawed premise the known rules missed?
Completer Does the implementation fulfil the promised capability?
Commission verification Does required external infrastructure actually exist and work?

Cure

Cure investigates the running system.

It scopes the application, builds coverage, dispatches specialized investigators, aggregates findings, manages fix work, retests corrected surfaces, reconciles the finding lifecycle and withholds certification when required evidence is incomplete.

Cure does not add unrequested features. It does not treat a model's summary as proof that persisted findings exist. It works from running behaviour and durable artifacts.

The contradiction that opened this document happened here, and the refusal to advance was this contract doing its job.

Principle: The artifact outranks the summary.

Caliber

Caliber is the codified structural quality system.

Its criteria represent known invariants across engineering, architecture, QA and process. Criteria can inspect:

Caliber is not an open-ended reviewer. It is strongest when the failure class is already understood well enough to encode.

What independent means here

FoundryOS uses independent narrowly and explicitly. The strongest current separation is:

This does not automatically establish a different model, different provider, outside organization or statistical independence. Those stronger forms remain evidence targets.

Codex audit

Codex review is used for open-ended technical and design judgment.

It can challenge:

The audit path is deliberately separate from the implementing context. Findings are folded into the work and the changed result is reviewed again.

For Auto Plan, pre-execution and post-execution Codex rounds are part of the profile. For non-Chassis work, go codex audit is the sanctioned entry.

Completer

Completer compares the promised system with the implemented system.

An application can pass runtime tests and still omit a required capability. Completer looks for that difference and produces a completion plan rather than silently reclassifying missing scope as future work.

Fresh context

Fresh context is used as an isolation boundary.

A reviewer that shares the builder's entire reasoning path can inherit the same assumptions. A fresh reviewer receives the relevant artifacts and authority without inheriting every narrative used to justify the implementation.

Fresh context is not automatically independent in a statistical sense. It is an architectural attempt to reduce correlated error.

Verifier disagreement

Different verifiers can disagree because they ask different questions.

A Caliber pass cannot erase a Cure runtime defect. A Cure pass cannot prove infrastructure outside the application is configured. A Codex audit can identify a novel concern that no existing criterion encodes. Completer can identify missing scope in an otherwise clean implementation.

FoundryOS should preserve the disagreement until the owning domain resolves it.

In the CrewOS build: the post-execution audit hit an ambiguous pre-merge checkout state and halted rather than passing it, and follow-on Cure work later exercised the running external-participant surface and returned a real finding.

9. RELEASE AND CONSEQUENCE CONTROL

Release and consequence control as four separate states: Close (prove the local terminal condition and record state), Promote (make the completed work reachable through the intended integration path), Prune (remove temporary work only after preservation is proven), Deploy (change the running system under consequence-appropriate authority). Cleanup is destructive work; it has to prove preservation first.

Contract: Completion, preservation, integration and deployment are separate states because each can fail independently.

Why this boundary exists. I lost work when cleanup followed partial success. Completion, preservation, integration and deployment can no longer collapse into one optimistic command.

The release sequence is:

This separation exists because a combined path caused real data loss.

Close

Close establishes that the local operation has reached its required terminal condition.

It collects evidence, verifies plan state, updates required paperwork and commits the coherent result. Closing does not automatically merge or deploy.

Promote

Promote makes the completed work reachable through the intended integration path.

It reconciles the target branch, submodule pointers, generated state and repository relationships. Promotion is about preservation and integration, not production deployment.

Prune

Prune removes worktrees, branches or temporary state only after reachability is proven.

The incident behind this gate is walked through the routing grammar in Section 3. The remedy was structural:

The lesson is simple:

Principle: Cleanup is destructive work. It has to prove preservation first.

Deploy

Deploy changes the running system.

It may include application deployment, live migrations, provider configuration, DNS changes and customer-visible consequences.

Deployment requires authority appropriate to the consequence. A successful merge does not authorize a production change by itself.

Rollback and destructive actions

Rollback paths need the same discipline as forward changes.

A destructive database correction, forced branch operation, provider deletion or customer-message action cannot be treated like an ordinary local edit. The system should identify irreversibility and require the appropriate human authority.

Terminal state

A legitimate terminal state is not "the model stopped."

It is a persisted state such as closed or elevated, produced through the execution contract and supported by the required evidence.

In the CrewOS build: the external-participant un-designation work landed at commit d20a60ea, and the post-execution lifecycle reached CLEAN at 93255817 through this controlled path.

10. MAINTENANCE AND INSTITUTIONAL LEARNING

The codification loop: an observed failure or operator catch leads to investigation, local correction, a generalized lesson, a Project Codex or Crucible update, a hook, test, schema, module or criterion, distribution, later verification, and retirement when warranted. Fix the failure, then fix the system that allowed it.

Contract: A preventable catch should become durable project or fleet knowledge, and where practical, a mechanism that makes recurrence harder.

Why this boundary exists. A patch closes one incident. Without codification, the next project pays for the same lesson again and the operator becomes the only durable memory. Every scar in this document ends up here, or it was wasted.

Caster

Caster maintains project and fleet knowledge.

It refreshes Project Codex material, reconciles documentation with source and observed state, collects reusable lessons, distributes relevant Crucible knowledge and records maintenance results.

Caster does not fix the application merely because it discovers stale knowledge. It returns the discrepancy to the owning engineering path.

Crucible

Crucible preserves reusable lessons across projects.

A lesson belongs in Crucible when it is sufficiently general, supported by evidence and relevant beyond the project where it was discovered.

Examples include:

Converge

Converge analyzes operator catches and repeated rescue work.

It asks:

Converge can surface whether a control may have become obsolete, but it does not currently own or generate a verified control-retirement proposal. Retirement remains an operator-and-plan action, and its governance is still listed as an open engineering problem.

The objective is earned autonomy, not indiscriminate automation.

A system earns more autonomy when repeated human intervention has been replaced with evidence-backed, bounded machinery.

The codification loop

A full example: CrewOS audit-chain integrity

During CrewOS Payment Plan Scheduler sandbox testing, audit-chain verification reported 20 mismatches and 1,142 broken links.

The cause was a migration that retroactively updated label fields on historical audit events.

The protected payload remained cryptographically intact because those labels were outside the hashed content. The immediate finding was therefore more subtle than "the audit chain was destroyed."

The dangerous pattern was historical audit-row mutation without a general rule describing what was allowed.

The response became fleet-wide:

  1. The incident was documented in CrewOS.
  2. The safe and unsafe mutation classes were separated.
  3. The Audit Chain Integrity Discipline was added to the Operating Standards.
  4. Caliber C49 was created.
  5. Fixtures covered allowed, forbidden, warned and grandfathered cases.
  6. Compatible projects in the current FoundryOS environment inherit the control.

A local failure became institutional knowledge and mechanical prevention.

Additional control mechanisms

The foundational article summarizes four failures for a business reader. The technical mechanisms are more specific:

Controls can also be removed

Improvement does not always mean adding another gate.

The Auto Chassis simplification retired criteria and machinery whose subject no longer existed. FoundryOS should remove controls when the underlying failure class has disappeared or a simpler mechanism now covers it.

Principle: A control earns its place by preventing a real failure class, not by having once been difficult to build.

Caliber keeps controls installed

When a control is important enough to become fleet doctrine, Caliber can protect the control itself.

This is how a lesson survives later refactoring. A hook, schema, route or release boundary can be tested not only for current behaviour but for continued presence and coherence.

In the CrewOS build: Caster refreshed the project knowledge after the work landed, and the build's audit-chain lesson is the origin of the fleet-wide C49 control described above.

11. CREWOS: FOUNDRYOS IN FLIGHT

A real CrewOS build trace in eleven steps: business capability, approved plan, durable planned execution on Chassis, parallel CrewOS worktrees, integration, source-ambiguity halt, corrected source conditions, post-execution audit, CLEAN, documentation refresh, follow-on Cure verification. 146 commits across multiple worktrees.

Contract: A system-in-flight case should show real components interacting without pretending every component participated.

The previous seven sections each ended inside this build. Here is the build whole.

The strongest current build example is the CrewOS external-participant capability work.

The Cast Run 61 build record reports 146 commits across multiple worktrees and a broad authorization surface. The figure is attributed to that operating record rather than presented as a reconstructable git-range count. The build exercised complex planning, durable execution, parallel construction, explicit integration, source-identity checks, post-execution review, documentation maintenance and follow-on Cure verification.

Authority and planning

The capability was represented as a plan rather than a collection of prompts. Scope, ownership, acceptance conditions, touched surfaces and verification requirements were explicit before broad implementation.

Durable execution

Persistent Chassis state carried the planned work across individual agent contexts while retaining stage, retry, decision and handoff information. The available build record supports durable planned execution on Chassis but does not directly identify the consumer profile by name.

Parallel construction and integration

Bounded work progressed across multiple worktrees. Integration remained a separate responsibility because clean branches can still disagree on authorization semantics, shared types and source state.

Ambiguity halted rather than becoming evidence

The audit encountered an ambiguous pre-merge checkout state and halted. It did not treat uncertain source identity as a clean verdict. After the source conditions were corrected, the external-participant un-designation work landed at commit d20a60ea, and the post-execution lifecycle reached CLEAN at 93255817.

Follow-on verification and maintenance

Caster refreshed the project knowledge after the build. Follow-on Cure work then exercised the running external-participant surface, including a finding in the External Participant Portal around the clear path. The construction result therefore remained subject to running-system verification rather than disappearing after merge.

Inspectable sources

The public article should point to exact sanitized evidence where available. Until the evidence package is published, the references below are classified as internal evidence only:

The case does not prove that FoundryOS can autonomously deliver every application. It demonstrates a substantial build crossing planning, execution, construction, integration, audit, maintenance and running-system verification.

12. CREWOS AND FOUNDRYOS GROWING TOGETHER

Eight stages of CrewOS and FoundryOS growing together: Valley Mainland production work, CrewOS conversion, Converter doctrine, larger CrewOS capability work, new verification and release failures, Caliber/Cure/Cell and operating controls, Auto Plan and Chassis for durable execution, and stronger future builds. Every stage made the next possible; the system closed the gaps at the system level, not only locally.

Contract: The architecture should explain the production history that forced it to take its current shape.

CrewOS was not produced by a complete FoundryOS system that already existed.

FoundryOS and CrewOS developed together. I did not design this system on a whiteboard. Production work forged it, one recorded failure at a time.

Valley Mainland began as a production application. Converting that application into CrewOS exposed problems that a greenfield demonstration would not have revealed:

Those problems produced reusable machinery and doctrine:

This history explains why FoundryOS has both specialized components and cross-cutting execution machinery.

The components own bounded responsibilities.

Chassis supplies reusable durable progression.

Auto Plan defines the cross-responsibility lifecycle for complex engineering plans.

DO NOT STOP and go codex audit cover smaller work that does not need the full lifecycle.

Cure uses Chassis because investigation and repair also need persistent state, retries, decisions and legitimate terminality.

13. OPERATIONAL FAILURE AND CONTROL MODEL

Ten failure classes mapped to primary control families: false completion (persistent plans, DO NOT STOP, Chassis terminality, Completer, post-execution audit); authority collapse (component boundaries, consequence classes, permissions, human gates, fail-toward-asking); context loss and drift (Project Codex, Required Reads, generated indexes, Caster, truth hierarchy); wrong observation point (source pins, visible root resolution, canonical bundles, scope checks); correlated verification (role, context and source separation, distinct verification domains); external-state mismatch (Commission preflights, live schema checks, provider inspection, running-system QA); unsafe release or cleanup (close/promote/prune/deploy separation, reachability checks, destructive-action gates); runaway or stalled automation (retry and cost bounds, novelty-aware convergence, persistent state, decision logs, elevation); procedural drift (Caliber criteria, generated projections, registry synchronization, fixtures, Caster); control-plane self-deception (claim-artifact checks, bundle provenance, Evidence Posture, re-derivable metrics). This is an operational reliability and control-plane model, not a complete adversarial security threat model.

Contract: The control plane must defend against failure modes created by capable agents, stale context, external systems and its own machinery.

This is primarily an operational reliability and control-plane model. It is not a complete adversarial security threat model. Repository prompt injection, secret exposure, malicious dependencies, poisoned fleet knowledge and compromised operator credentials require their own bounded security analysis.

14. GOVERNING THE CONTROL PLANE

Contract: FoundryOS machinery must earn its complexity by preventing consequential failure or reducing repeated operator work.

Deterministic where known, probabilistic where useful

FoundryOS divides work according to the nature of the decision.

Models and agents are used for:

Deterministic machinery is used for:

Humans retain authority for:

Operating Standards and Required Reads

Operating Standards hold binding cross-system doctrine.

Required Reads connect a task or component to the standards it must apply. This reduces the chance that an important rule exists in the repository but is absent from the active worker's context.

Hooks

Hooks enforce rules at lifecycle boundaries where failure becomes observable.

Examples include:

A hook should own a boundary, not become a hidden general-purpose policy engine.

Schemas, manifests and ledgers

Schemas make artifacts machine-checkable. Manifests declare scope, ownership and expected outputs. Ledgers preserve decisions, dispatches, retries, allocations and lifecycle state.

These mechanisms convert fragile prose expectations into inspectable contracts.

Registries and generated projections

FoundryOS uses registries for components, applications, criteria, monitoring surfaces and other controlled inventories.

Generated projections reduce hand-maintained duplication. The source registry remains authoritative, and drift checks ensure the generated view follows it.

Kill switches and safe degradation

Important automated mechanisms include explicit kill switches or manual fallbacks where appropriate.

The preferred degradation mode is not silent bypass. It is visible reduction to manual or operator-controlled execution.

Anti-bureaucracy test

A control should satisfy at least one of these conditions:

  1. It prevents a demonstrated consequential failure.
  2. It replaces repeated operator rescue.
  3. It protects a critical boundary or invariant.
  4. It produces evidence needed by a downstream decision.
  5. It allows a larger process to become simpler or safer.

A control that cannot meet that test should be simplified, retired or never added.

How to challenge a FoundryOS claim

A technical claim should be challengeable through five questions:

  1. Contract: What is the mechanism supposed to guarantee?
  2. Mechanism: Where is that guarantee implemented?
  3. Fixture: What test or counterexample proves it can fail?
  4. Run artifact: Where has it been exercised in operating work?
  5. Evidence class: Is the claim installed, enforced, exercised, outcome-evidenced, fleet-codified, comparative or portable?

This is more useful than arguing from component names or architecture diagrams alone.

15. OPEN ENGINEERING PROBLEMS

Ten open engineering problems and their next evidence targets: root-operator scalability (a second operator completes bounded consequential work without undocumented intervention); proportional governance (measure over- and under-governed tasks by work shape and consequence); control-plane economics (compare time, model use and prevented failure by control family); verifier independence (cross-model and cross-provider comparison on the same pinned evidence); control retirement (a retirement ledger with reason, replacement and post-removal observation); model and provider portability (run an equivalent lifecycle with a different model and transport); cross-stack generalization (exercise the same contracts on another application stack); measuring earned autonomy (track preventable intervention rate while preserving legitimate authority gates); evidence accessibility (publish sanitized, provenance-stable artifacts for major claims); business outcome evidence (define baseline, sample and measurement protocol before productivity claims).

Contract: Unresolved limits remain visible so the architecture does not convert ambition into evidence.

16. COMPONENT REFERENCE

Component Owns Primary return Important boundary
Charge Operational intake and discovery Structured business context Does not specify or build
Channel Canonical specification and authority Build-authoritative APP_SPEC and decisions Does not implement
Core Standard technical foundation Reusable scaffold, doctrine and modules Does not replace project-specific authority
Codex Template Project-memory structure Standard Codex skeleton Does not supply project facts
Crucible Reusable fleet knowledge Generalized lessons and patterns Writes are controlled through Caster
Conductor Lifecycle sequencing, bounded decomposition, coordination and integration readiness Bounded work packages and integrated execution state Does not perform all implementation or define business policy
Constructor Application construction inside approved scope Implemented application capability Does not define the cross-lifecycle route or invent policy
Completer Promise-to-implementation comparison Completion findings and plan Does not redefine the promise
Converter Existing-system conversion Converted system and reusable doctrine Not required for greenfield work
Cell Mobile and store readiness Mobile, PWA and distribution evidence Invoked only for registered targets
Commission Infrastructure provisioning Verified external configuration Repository state is not external proof
Cure Running-system QA and repair Findings, fixes, retests and certification state Does not add unrequested features
Caliber Codified structural governance Score, findings and criterion evidence Known rules, not open-ended judgment
Caster Project and fleet knowledge maintenance and distribution Refreshed Project Codex and distributed lessons Does not silently fix applications
Converge Operator-catch analysis and codification planning Codification proposal and surfaced control gap Does not silently implement recommendations or own control retirement

Shared machinery reference

Mechanism Role
Chassis Reusable durable execution framework
Auto Plan Complex engineering-plan profile on Chassis
DO NOT STOP Lightweight completion contract
go codex audit Canonical non-Chassis Codex-audit path
Project Codex Persistent project-specific memory
Crucible Fleet knowledge, subordinate to evidence and controlled distribution
Operating Standards Binding cross-system doctrine
Hooks Lifecycle-boundary enforcement
Schemas and manifests Machine-checkable artifact contracts
Close, promote, prune, deploy Separated release and consequence gates

17. GLOSSARY

C-component One of the 15 durable FoundryOS components that owns a bounded software-engineering responsibility.

APP_SPEC The canonical application specification produced through Channel. It is build authority, not merely descriptive documentation.

External Participant Portal (EPP) The CrewOS surface used to support participant access outside the primary internal-user model. EPP appears in some internal evidence paths and run artifacts.

Build authority The condition in which required business decisions, scope, ownership and evidence expectations are explicit enough for implementation to proceed without the builder inventing policy.

Auto Plan The Chassis consumer profile for complex planned engineering work. It carries work through draft, review, implementation, correction and closure.

Chassis The reusable, profile-driven execution framework that supplies persistent state, stage progression, decision routing, retries, handoffs and legitimate terminality.

Consumer profile The declarative data, seam wiring and procedure that tell Chassis what a particular lifecycle means.

Project Codex The maintained body of project-specific architecture, data, operational and workflow context required by future workers.

Crucible The fleet-wide store of generalized reusable engineering knowledge.

Codex audit Open-ended review of a plan or implementation through a role-, context- and source-separated Codex path. It is not automatically model-, provider- or organizationally independent.

go codex audit The sanctioned driver for Codex audits outside the Auto Plan lifecycle.

DO NOT STOP A lightweight execution contract that prevents silent premature termination and false completion on bounded work.

Source pin A recorded commit, branch, digest or source identity used to prove exactly what a verifier inspected.

Fold The act of applying review findings back into the plan or implementation before the next review cycle.

File bus One transport used to deliver canonical audit bundles and receive separate Codex review artifacts.

Chrome file-drop A second audit transport that uses the authenticated browser session when the file bus is unavailable or unsuitable.

Required Read A binding instruction that identifies the standards or project documents a worker must load before acting.

Two-signal surfacing The contract that combines session ownership with non-terminal Chassis state to prevent off-contract human surfacing during an active run.

Terminal state A persisted legitimate end condition such as closed or elevated, not merely the end of a model turn.

Keep-installed criterion A Caliber rule that protects the continued presence and coherence of an important control.

Root operator The accountable operator with cross-system routing and consequential authority. Dan Merchant currently holds this role.

18. SOURCES AND IMPLEMENTATION NOTES

This article is grounded in the current FoundryOS repository and its committed plans, standards, registries, evidence packages and run artifacts.

At this publication-candidate stage, repository paths and commit identifiers are internal evidence only unless a public or sanitized artifact is linked explicitly. They establish internal provenance, not public inspectability. The next evidence phase will classify each major claim as publicly inspectable, sanitized excerpt published, internal evidence only or public package forthcoming.

Key implementation sources include:

Incident provenance

Mutable metrics

Counts such as Caliber criteria, point scale, Crucible entries and engine size are implementation snapshots. They must be generated or re-verified at publication time. The August 2, 2026 values in this draft were rechecked against the current repository state.

The evidence package must publish the counting method beside each mutable metric. For Crucible, that includes entry-bearing files, exclusions, deduplication and contradiction handling. For Caliber, that includes shipped criteria, current scale, fixture coverage, blocking status, operating catches, retirements and known false-positive review limits.

Detailed Chassis reference figures are retained here rather than used as outcome claims: the current production engine source covers 25 modules and 3,743 lines excluding tests, while the published Cure v2 operation contains 49 ordered gates and 51 prerequisite edges. These figures describe implementation shape, not effectiveness.

19. CONCLUSION

FoundryOS does not ask one model to pretend it is an entire company. That was the first sentence worth saying in this document, and it is the last.

It engineers the responsibilities a software department must carry around capable models:

  1. Convert business reality into build authority.
  2. Preserve project context outside the session.
  3. Match execution weight to work shape and consequence.
  4. Build software and provision the external systems it depends on.
  5. Verify runtime, structure, promise and infrastructure through distinct paths.
  6. Separate completion, preservation, integration and deployment.
  7. Convert preventable catches into knowledge, controls or justified retirement.

The model provides intelligence.

The department turns that intelligence into a platform the business can depend on.

EVIDENCE

Every incident and mechanism in this article traces to a sanitized evidence package: the claim, the supporting artifact, what it does not prove, the redactions applied, and how to reproduce it.