Evidence
What controlled tests show, where Agent Enhancer helps, and what it costs.
Controlled benchmark · complete
93% fewer harmful outcomes in risky agent workflows
A harmful outcome is a duplicate mutation, a conflicting parallel action, or a provider rejection observed by the evaluator.
Two useful signals
Where the sidecar earned its place
Fewer parallel runs created duplicate or conflicting actions.
Runs completed with zero unnecessary sidecar calls.
Selective by design
Use it only when reliability work is justified
Parallel, repeated, scheduled, retryable, or duplicate-sensitive work.
One-time, read-only, low-risk tasks.
Trade-off
Reliability is not free.
Guarded workflows use more tokens and latency. Agent Enhancer therefore activates selectively, where preventing a duplicate or conflicting action is worth that overhead.
Evidence status
A completed benchmark, with real-world evidence still growing
Complete
Collecting
Current headline evidence comes from one synthetic benchmark and one model configuration. It is not a universal performance claim.
Technical proof and complete metricsMethod, distributions, corrections, diagnostics, and module health
Benchmark details
- Scenario acceptance
- 66/80 → 79/80
- Unresolved outcomes
- 0 → 0
- Valid Codex runs
- 200
- Infrastructure exclusions
- 4
Raw counters: duplicates 14 → 1, conflicts 12 → 1, provider rejections 0 → 0. The earlier 26 → 2 headline pooled correlated counters. The current 14 → 1 headline counts each affected run once.
Per-scenario distribution and cost
Paired p50 and p95 changes across 20 pairs per scenario. Positive values mean the guarded condition used more input tokens or time.
| Scenario | Affected runs | Unresolved | Input-token Δ p50 / p95 | Latency Δ p50 / p95 |
|---|---|---|---|---|
| Ambiguous create | 0/20 → 0/20 | 0 → 0 | +572.14% / +794.86% | +365.032% / +480.871% |
| Overlapping workers | 12/20 → 1/20 | 0 → 0 | +257.172% / +410.996% | +202.21% / +299.969% |
| Shared rate limit | 0/20 → 0/20 | 0 → 0 | +283.876% / +487.666% | +150.173% / +243.281% |
| Scheduled refresh | 2/20 → 0/20 | 0 → 0 | +240.176% / +365.178% | +139.603% / +294.024% |
| Low-risk control | 0/20 → 0/20 | 0 → 0 | -33.913% / +1.166% | -32.53% / +3.312% |
This distribution is a post-hoc descriptive reanalysis of unchanged frozen rows. Low-risk negative deltas are host variance, not product savings: the sidecar made zero calls in those runs.
Targeted v1.7.2 diagnostic
All holder claims were prepared before the write.
Exactly one mutation in every accepted trial.
Separate exploratory diagnostic, not pooled with the 200-run publication.
Fixtures, exclusions, and proof layers
Four infrastructure exclusions were rerun unchanged. The condition-blind evaluator was invoked for every risk run.
Synthetic fixtures preserve exact injected failures and expected machine state.
The negative result led to the skills-first, on-demand design.
Production module evidence
- Current self-tests
- 24 / 24
- Direct actions checked
- 74
- Current datasets
- 2
- Usage evidence
- publishing
| Module | Mode | Self-test | Latency | Up to 30d completion / sample |
|---|---|---|---|---|
| PennyLock 0.2.0 | atomic | ok · current | 39 ms | < 10 observations |
| Global Seen Stamp 0.2.0 | atomic | ok · current | 44 ms | < 10 observations |
| Exactly-Once Baton 0.2.0 | atomic | ok · current | 44 ms | < 10 observations |
| Negative Cache Ticket 0.2.0 | atomic | ok · current | 43 ms | < 10 observations |
| Status-Code Forge 0.1.0 | ephemeral | ok · current | 45 ms | < 10 observations |
| Safe Synthetic Fixture Vault 0.1.0 | lookup | ok · current | 1 ms | < 10 observations |
| Error-Code Cemetery 0.1.0 | lookup | ok · current | 1 ms | < 10 observations |
| 402 Error Rosetta Stone 0.1.0 | lookup | ok · current | 1 ms | < 10 observations |
| Worked-Once Recipe Vault 0.1.0 | lookup | ok · current | 1 ms | < 10 observations |
| Swarm Semaphore 0.2.0 | atomic | ok · current | 42 ms | < 10 observations |
| Barrier Bell 0.2.0 | atomic | ok · current | 41 ms | < 10 observations |
| Freshness Lease 0.2.0 | atomic | ok · current | 41 ms | < 10 observations |
| Failure Sequence Forge 0.1.0 | ephemeral | ok · current | 44 ms | < 10 observations |
| Swarm Rate Gate 0.2.0 | atomic | ok · current | 45 ms | < 10 observations |
| MCP Tool Contract Linter 0.1.0 | lookup | ok · current | 3 ms | < 10 observations |
| x402 Requirement Drift Diff 0.1.0 | lookup | ok · current | 2 ms | < 10 observations |
| Webhook Attempt Meter 0.1.0 | ephemeral | ok · current | 40 ms | < 10 observations |
| MCP Capability Handshake Diff 0.1.0 | lookup | ok · current | 1 ms | < 10 observations |
| x402 Facilitator Compatibility Diff 0.1.0 | lookup | ok · current | 1 ms | < 10 observations |
| MCP Elicitation Safety Linter 0.1.0 | lookup | ok · current | 1 ms | < 10 observations |
| x402 Quote Fingerprint Guard 0.2.0 | atomic | ok · current | 44 ms | < 10 observations |
| MCP Schema Edge-Case Atlas 0.1.0 | lookup | ok · current | 1 ms | < 10 observations |
| Workflow Guard Planner 1.1.0 | lookup | ok · current | 2 ms | < 10 observations |
| Opaque Workflow Checkpoint 1.2.0 | atomic | ok · current | 42 ms | 100% / 12 |