forge-harness: Engineering Methods for Robust AI Collaboration Harnesses
Authors/Creators
Description
VERSION 1.0.2 — CLAIM-WITHDRAWAL RELEASE (2026-09-19). This version withdraws four claims that v1.0.1 stated more strongly than its design supports: a 100% three-model agreement rate (one run per model, no repeated runs, no pre-registered rubric); an equivalence claim about context-isolation level, for which no non-inferiority margin was pre-specified; and the phantom-rate (<10%) and reachability (≥80%) figures, which are restated as observed working values rather than deployment gates. Independent Architectural Convergence moves from §4.6 to §2.6 and is marked as literature context rather than evidence. A scope paragraph is added to §1 stating that forge-harness names the harness layer and that pipelines assembled from it are instances of that layer. The v1.0.1 corrective-release note is retained below.
VERSION 1.0.1 — CORRECTIVE RELEASE (2026-09-06). This version corrects systematically incorrect bibliographic metadata in v1.0. Every arXiv identifier cited in v1.0 resolves to a real paper, but for eleven of them the recorded title, authors, or institution belonged to a different paper — in one case to a paper in an unrelated field. All references have been re-verified against the arXiv API (21/21 titles and first authors now match), and every load-bearing external numeric claim has been re-checked against the cited paper's full text.
Substantive consequences, itemised in the Erratum record in Appendix E of the document (a short erratum notice follows the abstract): the 98.4% harness-infrastructure figure is restated as a codebase-composition ratio reported as a community estimate, not a measurement of session content, and the accompanying "10,000 Claude Code sessions" claim is removed as unsupported; AHE's 11.8% and 33.7% are restated as regression-precision and fix-precision against their random baselines; a claim attributed to arXiv:2605.28065 is removed because that paper does not make it; Ptah's verifier axes are corrected; and the "16 parallel agents validated in production" attribution to Anthropic's Dynamic Workflows is withdrawn — Anthropic describes a research preview reporting hundreds of parallel subagents and states no 16-agent figure, so the tier boundary is now stated as a design choice rather than an externally validated threshold.
No measurement in Sections 4.1–4.5 or Table 1 depends on the corrected references; those are measurements of the toolkit itself and are unchanged. The eleven-implementation convergence count is unchanged — all eleven exist; what was wrong was how four of them were named. Version 1.0 remains published as the record of what was distributed.
This paper presents four complementary harness engineering methods addressing the primary failure modes in AI collaboration harnesses: steel-quench (adversarial structural validation), source-grounding-audit (phantom claim
detection), harvest-loop (session-to-harness self-evolution), and sim-conductor (pre-deployment transfer validation). Applied to forge-harness itself: 10 structural defects resolved (4 S-grade, 4 A-grade), phantom claim rate reduced from 6.4% to 0% (3/47 → 0/44), 100% skill reachability confirmed across 4 external personas, 80% HIGH-grade external contribution absorption rate.
Files
forge_harness_v1.0.2.pdf
Files
(926.2 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:c38ef882587adfafe6d56e04d92a9035
|
926.2 kB | Preview Download |
Additional details
Dates
- Updated
-
2026-05-29Adds prompt-regression, mcp-circuit-breaker, and token-budget-gate skill domains. Expands independent architectural convergence evidence from 3 to 6 implementations (SwarmHarness, HarnessAPI, Harness-Bench). Updates skill count to 28.
- Updated
-
2026-05-30Nine independent implementations (up from six) converge on the same outer-loop architecture — SkillOpt (arXiv:2605.23904), AHE (arXiv:2604.25850), and Scaling the Harness (arXiv:2605.26112) each independently formalized patterns forge-harness had already operationalized (synthesizer gate, regression blindspot, stale-but-confident detection). Three new gap skills added from PR #32: edit-manifest (prediction-verification loop), memory-hygiene (stale-but-confident detection), VCS-Layer Gate Enforcement (git pre-commit hook + marker file pattern). Skill count: 30 (26 fh-meta + 4 fh-commons). Quantitative summary table consolidated. Explicit positioning vs. performance optimization systems.
- Updated
-
2026-05-30Eleven independent implementations converge on the same outer-loop architecture — two new convergence points: Ptah (arXiv:2605.29861, stage-wise multi-agent verification convergent with 3-axis auto-gate) and Anthropic Dynamic Workflow (parallel sub-agent orchestration at scale, convergent with agent-composer Wave architecture). sim-conductor updated to task-adaptive persona selection (3-tier sourcing: installed plugins → built-in role directives → external fetch; scale 3–16 parallel agents). Model-agnostic harness layer positioning added (§5.4): Base mode (Sonnet) sufficient for standard validation; Amplified mode (Opus orchestrator + Sonnet executors) extends to Dynamic Workflow-scale fan-out without changing the validation contract. Table 1 convergence count corrected to 11.
- Updated
-
2026-05-30Two new convergence points (total: 11): Ptah [arXiv:2605.29861] — stage-wise multi-agent verification convergent with 3-axis auto-gate; Anthropic Dynamic Workflow — parallel sub-agent orchestration at scale, convergent with agent-composer Wave architecture. sim-conductor: task-adaptive persona selection (installed plugins → built-in fallback → external fetch), scale 3–16 parallel agents. Model-agnostic positioning (§5.4): Base (Sonnet) for standard validation; Amplified (Opus orchestrator + Sonnet executors) for Dynamic Workflow-scale fan-out — same validation contract either way. Pipeline architecture diagrams added (Figures 1–2).
- Updated
-
2026-09-06Corrective release v1.0.1. Eleven arXiv references carried the title, authors, or institution of a different paper — one of them a paper in an unrelated field. All references re-verified against the arXiv API (21/21 titles and first authors match). Substantive consequences: the 98.4% figure restated as a codebase-composition ratio reported as a community estimate, and the accompanying "10,000 sessions" claim removed as unsupported; AHE 11.8%/33.7% restated as precision against their random baselines; a claim attributed to arXiv:2605.28065 removed because that paper does not make it; the "16 parallel agents validated in production" attribution to Anthropic withdrawn. No measurement in Sections 4.1-4.5 or Table 1 is affected.
- Updated
-
2026-09-07Erratum relocated from the front matter to Appendix E, leaving a short erratum notice after the abstract; a first-time reader now reaches Section 1 without reading two pages of corrections first. Record metadata brought in line with the corrected paper: the References field still carried all eleven pre-correction attributions and was replaced with the verified list (22 to 24 entries).
Software
- Repository URL
- https://github.com/chrono-meta/forge-harness
- Development Status
- Active
References
- Liu, J. et al. "Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems." arXiv:2604.14228, 2026.
- Ning, X. et al. "Code as Agent Harness." arXiv:2605.18747, 2026.
- Seong, H. et al. "The Last Harness You'll Ever Build." arXiv:2604.21003, 2026.
- Lee, Y. et al. "Meta-Harness: End-to-End Optimization of Model Harnesses." arXiv:2603.28052, 2026.
- Sengupta, S. et al. "Meta-Engineering Harnesses for AI-Native Software Production: A Contract-Driven Adversarial Verification Architecture with Early Deployment Report." arXiv:2605.25665, 2026.
- Zhu, J. et al. "Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems." arXiv:2605.26302, 2026.
- Benkovich, N. et al. "Agyn: An Open-Source Platform for AI Agents with Scalable On-Demand Execution, Agent Definition as a Code, and Zero-Trust Access." arXiv:2605.27575, 2026.
- Diks, I. et al. "Verifiable Benchmarking of Long-Horizon Spatial Biology." arXiv:2605.28065, 2026.
- Jose, E. "SwarmHarness: Skill-Based Task Routing via Decentralized Incentive-Aligned AI Agent Networks." arXiv:2605.28764, 2026.
- Jose, E. "HarnessAPI: A Skill-First Framework for Unified Streaming APIs and MCP Tools." arXiv:2605.22733, 2026.
- Yao, Y. et al. "Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows." arXiv:2605.27922, 2026.
- Yang, Y. et al. "SkillOpt: Executive Strategy for Self-Evolving Agent Skills." arXiv:2605.23904, 2026.
- Lin, J. et al. "Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses." arXiv:2604.25850, 2026.
- Gu, S. "From Model Scaling to System Scaling: Scaling the Harness in Agentic AI." arXiv:2605.26112, 2026.
- Zhang, C. et al. "Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation." arXiv:2605.29861, 2026.
- Anthropic. "Introducing Claude Opus 4.8" (dynamic workflows, research preview for Claude Code Enterprise / Team / Max). https://www.anthropic.com/news/claude-opus-4-8, 28 May 2026.
- Christi, R. harness-evolver. MIT License. https://github.com/raphaelchristi/harness-evolver, 2026.
- Kwon, S. forge-harness. https://github.com/chrono-meta/forge-harness · DOI 10.5281/zenodo.20397566, 2026.
- Perez, E. et al. "Red Teaming Language Models with Language Models." arXiv:2202.03286, 2022.
- Bai, Y. et al. "Constitutional AI: Harmlessness from AI Feedback." arXiv:2212.08073, 2022.
- Maynez, J. et al. "On Faithfulness and Factuality in Abstractive Summarization." ACL 2020. arXiv:2005.00661.
- Ji, Z. et al. "Survey of Hallucination in Natural Language Generation." ACM Computing Surveys 55(12), 2023. arXiv:2202.03629.
- Shinn, N. et al. "Reflexion: Language Agents with Verbal Reinforcement Learning." NeurIPS 2023. arXiv:2303.11366.
- Madaan, A. et al. "Self-Refine: Iterative Refinement with Self-Feedback." NeurIPS 2023. arXiv:2303.17651.