CFC Cross-Model Benchmark v1: Claude and Gemini Evaluation (600 Primary Runs)
Authors/Creators
Description
I’m sharing the frozen CFC Cross-Model Benchmark v1 evaluation covering 600 primary runs: 300 Claude runs and 300 Gemini runs across the same 100 decision-closure variants.
Both model environments achieved the same strict semantic PASS rate: 296/300 (98.67%). However, the observed semantic anomalies occurred on different variants and through different mechanisms. Claude’s failures included reasoning attribution, stale-state carryover, instruction-recognition failure, and false closure from duplicate priority coverage. Gemini’s failures were dominated by state/layer binding errors.
The anomaly variants did not overlap in these frozen runs. I treat that as a descriptive observation, not as evidence of systematic model complementarity.
The package also preserves the provenance limitation for Gemini CM-V002 and CM-V003. Excluding all six associated runs gives 290/294 semantic PASS (98.64%), so the aggregate conclusion is materially unchanged.
The narrow conclusion is that a near-99% semantic PASS rate on this benchmark does not eliminate rare state-transition, state-binding, instruction-following, and decision-closure failures.
This does not demonstrate that CFC prevents these errors. That requires a separate controlled comparison using the same frozen benchmark with and without the CFC/controller layer.
I’m sharing the package as a baseline reliability record and as groundwork for that next intervention study.
Where preserved and provenance-verified, the full benchmark questions/prompts and corresponding Claude and Gemini responses may be made available on request for independent inspection, scoring review, or methodological audit.
Any such release would preserve the original frozen materials and would clearly identify any items for which source provenance could not be independently verified.
Files
CFC_Claude_vs_Gemini_600_Primary_Runs_Comparison.md
Files
(116.6 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:9997b637adc83ab89d409b45d7f21690
|
11.6 kB | Preview Download |
|
md5:0fbd990413cfb4e0df7aa07f06b58cb2
|
43.0 kB | Preview Download |
|
md5:446ec266ba29dc1b1c8bf12e55033123
|
45.2 kB | Preview Download |
|
md5:02094b741ae4d4f41aee7725a6cf95cc
|
383 Bytes | Preview Download |
|
md5:766b078a5e36d2f5950be59a4c234770
|
1.4 kB | Preview Download |
|
md5:9c926d6bed68884a21951ba6edc5ccc7
|
353 Bytes | Preview Download |
|
md5:08cbe1db0a483d06ca6863c2b06b6831
|
3.9 kB | Preview Download |
|
md5:080b620a5445b55d255965df6d0b88d9
|
2.7 kB | Preview Download |
|
md5:1d1e2e8c86a19be6f859bee3cda4a603
|
2.0 kB | Preview Download |
|
md5:128fc7b4d021615db87a18bafe52b300
|
1.1 kB | Preview Download |
|
md5:261dce83cfd5b7da1a03c41e8bb7418b
|
4.8 kB | Preview Download |