I kept running into the same failure mode in agent systems: a tool call succeeds, so the agent believes the goal was met; a value is retrieved, so it's treated as true; an LLM says "PASS", so that verdict quietly becomes evidence for the next step. Each step looks fine. The composition is what breaks.
Ario is an attempt to make that composition step mechanically checkable, not just promptable.
What's actually implemented
core/schema.py + core/audit_engine.py: a deterministic audit over typed artifacts (Claims, HistoricalStates, LineageEdges, RetrievalEvents, Evidence, Assessments, AuditResults). Inputs are JSON. The engine never touches the clock, the network, or the filesystem.
ario_agent.py: a bounded workspace agent. Tasks are declarative JSON. Tools are allowlisted — read_text, file_fingerprint, git_status, compile_python, run_tests, replace_text, restore_backup, verify_text, recovery_preflight, inspect_backup. No arbitrary shell, no network, no delete, no privilege escalation.
replace_text requires a SHA-256 precondition, writes a byte-for-byte backup outside the workspace, uses os.replace for atomicity, and verifies the post-write hash. It does not auto-rollback if a later verification fails; rollback is a separate guarded task. That limitation is intentional and stays visible in the ledger.
ario_workflow.py: declarative workflows, max 8 predeclared stages, exact status-based branching (when: {stage_id, status}), no expression evaluation, no dynamic planning.
ario_planner.py: optionally asks a local Ollama model for a recommendation — one implementation/test pair and one engineering question. The model cannot emit workflow stages, action IDs, tools, conditions, or writes. Ario builds the workflow deterministically afterward. Execution is opt-in.
Append-only JSONL ledgers, exclusive lock files, duplicate-task-ID refusal, malformed-ledger fail-closed, pytest regression suite covering F05/F07/F09/F11/F15 adversarial fixtures.
The design decisions I care about
Verdicts cannot become evidence. AuditResult is a distinct artifact kind from Evidence. You cannot satisfy Assessment.admissible_evidence_refs with an audit_id, even if you set reference_type = "EVIDENCE". The engine inspects the actual supplied collections, not the caller's label. If an ID resolves across kinds or has duplicates: UNKNOWN, not a preferred kind.
Local PASS does not compose into global truth. F15 is a typed CompositionRequest with explicit participants, a versioned rule, and a typed requested_conclusion. Requesting GLOBAL_TRUTH, CLAIM_TRUTH, SAME_ENTITY, ONTOLOGICAL_IDENTITY, or CONSCIOUSNESS returns COMPOSITION_FORBIDDEN. BOUNDED_SUMMARY passes. Enforced structurally, not by prompt.
Rule-version gating. F05/F07/F11/F15 classifications only activate if the matching M0-*-1.0 is declared. A later rule cannot retroactively change what an earlier invocation meant.
Fail-closed everywhere. Malformed ledger: refuse. Duplicate task_id: refuse. Stale lock: report, don't auto-remove. Missing precondition hash: refuse. Ambiguous ID: UNKNOWN.
What is deliberately not implemented
No LLM semantic judge in the audit path.
No ontological claim. self-claim ≠ self-evidence, memory ≠ truth, retrieval success ≠ source continuity.
No production hardening. Ledgers are local files. A local admin can modify them. Cryptographic chain integrity, OS isolation, and approval workflows are listed as future work.
The F15 firewall covers named escalation targets; it does not cover every possible laundering route.
Why
Most agent failures I've seen are epistemic, not functional. Ario doesn't prove the system correct. It makes the composition step checkable, and prevents silent self-upgrade.
Repo: https://github.com/SepehrGhanbari/Ario
Critique welcome. Three things I'm least sure of:
Whether the audit surface generalizes beyond fixtures.
Whether "verdict cannot become evidence" survives a downstream consumer that relabels output.
Whether append-only JSONL is useful or theater until there's a cryptographic chain.