Building Decision Firewall, then narrowing it
A refund workflow grew into a framework. Testing it against a competent application changed the scope I kept and the claims I was willing to make.
I started Decision Firewall with a refund request. A model could read the customer's message, but I did not want its interpretation to become permission to move money. I wanted the payment record to establish the facts, and explicit policy to decide what could happen.
That was a useful problem. The project that grew around it was harder to justify.
By the time I asked for an independent comparison, I had a reusable runtime, historical-data workflows and an evaluation layer. The comparison found equal correctness against competent application controls, more integration friction, and a slower implementation. I kept the reference implementation and stopped the adoption claims.
This is how I got there, including the parts I would now do in a different order. I directed the work through coding agents, questioned the measurements and commissioned a separate agent evaluation. I did not personally hand-write every implementation, and the evaluation did not measure human productivity.
A refund is small enough to explain
I picked subscription refunds because an incorrect action has a visible consequence. A duplicate-charge claim needs transaction evidence. A cancellation may depend on when it happened and whether the service was used. A previous partial refund changes the remaining balance.
The demonstration policy made those distinctions concrete. Verified duplicate charges could be refunded against the duplicate payment. Cancellation qualified within seven days only if the subscription was unused. Eligible refunds up to ₹1,000 could proceed automatically; larger ones needed review. A configurable ₹10,000 daily automated limit routed excess demand to review. These were example rules, not statements of legal entitlement. Every payment effect was simulated. Original demonstration policy.
I also kept customer value and retention pressure out of eligibility. A convincing message about a valuable customer must not manufacture a duplicate payment. That separation mattered more than the model's confidence score.
- 1An assessment is an interpretation. The message sounds like a duplicate-charge complaint.
- 2An authorization is permission for an exact action. Refund this amount against this payment, using these facts and this policy, before this permission expires.
Every next layer sounded reasonable
I did not want to stop at a refund application. I asked for a framework that other developers could use with their own decision models and business rules. Refunds became a domain pack; access control and a simulated deployment example tested the boundary around the core.
Then I asked what happened before and after the decision. How would someone import previous decisions? Where would reviewed labels live? Could they compare policies without executing anything? Could telemetry explain what happened? I had LangChain's abstraction style and Langfuse's observability experience in mind, without a reason to reproduce either product.
The plan acquired CSV/JSONL history ingestion, immutable dataset snapshots, optional historical retrieval, named Python rules and offline experiments. Those features had sensible constraints: a historical approval must never become a live authorization; held-out examples must not leak into model context; a policy experiment must never invoke an executor.
I later added preparation before inference. It could validate structured intent and request missing evidence before spending a model call. It still needed final authorization against fresh facts. The extra stage was a capability to test, not an automatic improvement.
The mistake was treating a coherent feature list as enough evidence for the package around it.
The laptop put useful limits on the design
My local machine had a GTX 1650 with 4 GB VRAM and roughly 8 GB system RAM. I chose Laya as the first decision-model adapter and left Jev out of this release. The fixture path had to work without downloading any model.
The recorded local run used convaiinnovations/laya-typed-decisions, pinned to revision 1a793eb568e6718f15941d08f85432581df534e3, with SDK 0.3.20. It ran on CUDA in FP32. Cold loading took 28.42 seconds; median recorded inference was 47.45 milliseconds across 34 assessments, including the first inference. Those are local measurements, not an accuracy result. Local-model report.
On a narrow screen, scroll the table sideways.
| Part | Choice and responsibility |
|---|---|
| Core contracts | Python 3.12 and Pydantic for validated proposals, assessments and lifecycle records. |
| Durable state | SQLite transactions for authorization claims and reservations; a separate simulated effect ledger with idempotency keys. |
| Audit | RFC 8785 canonical JSON, SHA-256 and Ed25519 signatures for linked records and verifiable receipts. |
| Developer surface | Typer SDK/CLI workflows. The optional refund inspector uses FastAPI and Jinja templates. |
| Model and telemetry adapters | Optional Laya/PyTorch; optional OpenTelemetry export through OTLP. Neither is required by the model-free core. |
I kept model output away from execution credentials. Domain packs supplied facts and policy; executors owned the downstream effect boundary. A signed record can prove that recorded content has not changed under a trusted key. It cannot prove that the evidence was true, or protect a machine whose signing key has been compromised. Architecture and trust boundaries.
Compared with what?
“Is the same foot just sitting at two different places, or genuinely there is some distinction?”
I had asked for with-and-without numbers. The important part was deciding what “without” meant. A model classification mapped straight to a refund is an easy opponent. A competent application already checks the ledger, binds approval to an action and handles uncertain execution.
The synthetic refund comparison therefore had a deliberately minimal model-directed arm and two governed arms. Both governed approaches saw the same evidence and fault schedules. Both used the same simulator, including its own balance and destination protections.
In the controlled-fixture condition, application controls and Decision Firewall each produced zero contract-violating effects and zero unresolved attempts after recovery. The minimal arm produced 95 violating effects across 170 request trials. That demonstrated why application controls matter. It did not establish a special property of the framework.
The framework also cost time. On the matched straightforward-success subset, the application's p95 was 26.52 ms and the framework's was 67.89 ms. The paired median added cost was 23.23 ms across 25 paired trials. The full run had 32 frozen scenarios repeated five times; repetition measured runtime variation, not 160 independent business situations. Protocol and measured refund results.
I had also questioned an increasing “p value.” Here the reported p50 and p95 were latency percentiles, not statistical p-values. Higher p95 meant a slower tail in that run. Calling it a percentile did not make the overhead disappear.
A classifier cannot demonstrate authorization value
I asked to run the full available data rather than rely on a convenient sample. Across BANKING77 and Bitext, the comparison accounted for 39,955 source rows. The classifier was CPU naive Bayes, not Laya. Wrapping an unchanged classifier did not change its predictions.
On a narrow screen, scroll the table sideways.
| Evaluation | Rows | Accuracy in every arm |
|---|---|---|
| BANKING77 official test | 3,080 | 78.60% |
| BANKING77 training rows, out-of-fold | 10,003 | 78.10% |
| Bitext, out-of-fold | 26,872 | 98.31% |
I kept the official test separate from the out-of-fold scores. Bitext's split was ours, not a publisher leaderboard split. The prepared BANKING77 test path added a paired median 13.527 ms, with no prediction changes. There was no action to authorize in this task. Unless I independently needed validation or signed preparation records, I had added work without helping classification. Full-data report and split limitations.
I then asked for tool-using tasks. Laya could not fill the conversational agent and simulated-user roles, so a later development pilot used openai/gpt-oss-20b through OpenRouter, with a token cap, seeded requests and bounded rate-limit backoff. It scheduled 90 episodes across selected τ-bench and AgentDojo tasks; 78 completed and 12 remained incomplete. This was not a full-suite score.
On the attacked AgentDojo subset, the baseline recorded 8 attacker successes in 9 completed episodes; application controls and the framework each recorded 0 in 10. Both controlled arms completed the user task in only 4 of those 10 episodes. Blocking an attack while failing a legitimate request was part of the result. The controls matched again. The model judge and selected attack also limit what I can infer. Later agent pilot.
Audit was useful. Telemetry was not a free win.
I thought observability might be the stronger reason to reuse the framework. An application team would otherwise write lifecycle instrumentation and decide how to connect an execution attempt to the evidence behind it.
The local reuse study made that concrete. The application had eight explicit event call sites and four operation decorators; the framework integration passed an observer and needed no application event hooks. Framework receipts verified and rejected tampering in 60 of 60 synthetic cases. The application baseline had unsigned logs, so signing was additional bundled functionality, not a capability it was unable to implement.
The same report showed gaps. In the measured callable-policy path, framework spans lacked the evaluation ID, policy version, disposition and reasons. The manually instrumented application exported the first three; the shared allowlist omitted reasons in both. Neither exported any of the four sensitive canaries. Both completed an effect despite exporter failure, while required local audit failure blocked execution. Observability measurements.
I now separate those contracts. Telemetry helps me find the relevant records and may be dropped under backpressure. The required audit path participates in the authorization contract. A dashboard cannot substitute for a durable answer about whether an action happened.
That distinction is worth keeping. It still does not establish that a team saves investigation time by adopting this runtime. Telemetry alone is a weak reason to accept the whole dependency when its observer utility can be reused separately.
The comparison that changed the recommendation
I had no human participants available, so I commissioned a separate Claude Code agent evaluation. One builder used the framework; another built equivalent application controls against the same purchase-order requirements. Policy maintenance and incident investigation followed.
This was independent of the framework builder, not independent in the sense of a large external study. There was one task and one builder per arm, using the same model. I wanted an adversarial check on the project, not an agent-generated testimonial.
On a narrow screen, scroll the table sideways.
| Measured item | Framework | Application |
|---|---|---|
| Visible acceptance tests | 26/26 | 26/26 |
| Held-out acceptance tests | 114/115 | 114/115 |
| Policy-change tests | 15/15 | 15/15 |
| Investigation score | 48/48 | 48/48 |
| Build tool calls | 45 | 34 |
| Policy-change tool calls | 32 | 20 |
| Library-contract workarounds | 8 | 0 |
| Held-out suite time | 39.6 s | 4.5 s |
The application also recorded two test/path workarounds; zero above refers specifically to library-contract workarounds. Both arms failed the same approval-after-outage test. The frozen spec was ambiguous about invalidating the authorization versus invalidating the approval as well. No downstream effect occurred. I kept the scored failure rather than rewriting the result.
The framework builder requested seven core changes and implemented eight workarounds. Its implementation duplicated many issue-time checks to get specific refusal reasons before calling the framework. At that point reuse was visibly fighting the integration.
There were important confounds. The 374-line requirements already supplied much of the lifecycle design. Investigators received downstream exports, and neither could run its operator CLI because of a harness import problem. The framework builder hit a spend limit and was resumed. Held-out tests were withheld, not access-controlled; the evaluator's tool-call audit found no builder access. The report records the remaining deviations.
I cannot turn those agent tool counts into developer hours. I also cannot use the limitations to claim a benefit that was not measured. Independent comparison and full caveats.
“Stop feature expansion and adoption/migration claims.”
I fixed the defects without rewriting the study
The review produced 21 findings, with none rated critical or high. Four were confirmed defects. An approval could bind to evidence refreshed after the evaluation the reviewer saw. Deactivating a rejecting reviewer could revive an older approval. An oversized integer in adapter metadata could strand an action that had already happened. And the repository tests failed to detect six of nine injected runtime faults.
That last result made the green suite less reassuring. Some downstream checks were hiding missing guarantees in the layer I was trying to evaluate.
The 0.3.1 correction required the inspected evaluation ID when approving and rejected material changes. It preserved the latest rejection for that proposal revision. It encoded large metadata integers losslessly for canonical audit and added tests aimed directly at the core invariants.
The recorded Windows verification finished with 298 passing tests, including 22 corrective tests, and detected all nine targeted mutations. That is bounded regression evidence. It is not exhaustive fault coverage, and it is not a rerun of the adoption pilot. Old reports remain unchanged. Corrective release and migration boundary.
A useful control can still be an expensive abstraction
The global SQLite write lock remains. It spans adapter I/O and explains the roughly ninefold held-out-suite time in the pilot. That ratio describes that suite, not every request or deployment. Releasing the lock safely would require a different concurrency and recovery design; it is not a cosmetic performance patch.
I did not need another architecture project to avoid saying that this implementation was slow. I needed to document its operating limit and narrow the recommendation.
I also overreached in the sequence of work. History ingestion and model-context budgets can each be useful, but I built that surrounding workflow before establishing whether an unfamiliar integrator benefited from the central runtime. Classification benchmarks answered a question the framework was not designed to win. A refund-specific browser added a second lifecycle surface to maintain.
Some controls remain necessary: exact action binding, unresolved-outcome recovery and authoritative downstream idempotency. The packaging around them has to earn its place. Moving application checks into a reusable dependency does not make those checks more correct merely because the dependency has a name.
The review also found genuine limits to generality: an assessment/message-shaped submission contract, one caller identity per runtime instance and a fixed approval lifetime. A domain-neutral core can still impose awkward assumptions on another domain.
A reference implementation is a legitimate outcome
The core claims did not all collapse. The independent review verified action/evidence binding and recovery without replacement dispatch. A capacity-one resource admitted exactly one dispatch across eight operating-system processes. Historical replay called neither the evidence resolver nor the executor.
I kept that implementation and the conformance kit. They give a builder concrete failure cases to inspect, including stale evidence and a lost response after success. I am comfortable offering those artifacts for study. I am not recommending that an application with equivalent controls migrate to obtain an unmeasured advantage.
The kit also makes the distinction testable outside the model. This is the model-free entry point in the repository; use Python 3.12 or newer and activate the virtual environment for the local shell before installing:
git clone https://github.com/asadani/decision-firewall.git
cd decision-firewall
python -m venv .venv
# Activate .venv for your shell, then:
python -m pip install .
firewall framework --domain refunds demo
firewall conformance decision_firewall.conformance_fixtures:deployments --save conformance.jsonThe deployment fixture covers 17 declared checks. Unsupported cases are reported separately; a pass is not production certification. Authentication, credential isolation and the authoritative effect boundary still belong to the integrating deployment. Conformance contract.
Write the comparison before the expansion
My useful contribution was steering the work toward a harder question. I asked whether the baseline was competent, whether all rows were counted, whether a blocked call also blocked a legitimate task, and whether supplied telemetry actually exported the fields I needed. I should have asked the first question before authorizing the wider framework plan.
On the next build, I would implement one consequential workflow and its modular application comparator before adding the developer platform around it. I would decide what would count as a reason to adopt: fewer integration defects, less maintenance effort, or faster incident reconstruction. Then I would choose a study that can measure that reason.
A future test could use several integrations, less prescriptive requirements and incidents without an answer-bearing downstream export. Human participants would be needed for human-productivity claims. That is a possible follow-up, not an explanation for dismissing the result I have.
I started by trying to make model-driven actions accountable. I ended by having to make the framework's own claims accountable. Keeping the code was easy. Narrowing what I said about it was the more useful engineering decision.