Published August 7, 2026 | Version v1

Cross-Model Critical Audit Evaluation: Gemini vs Chmurka — A 12-Run Pilot Study

Authors/Creators

Abstract (Polish)

Twelve valid evaluation runs were conducted using Gemini and Chmurka across two conditions: neutral critical review and an academic/professor-style review.

The evaluation focused not only on numerical ratings, but also on observable behaviours, including:

  • false positives or unsupported confirmations,
  • important omissions,
  • instruction compliance,
  • evaluation of proposed corrections,
  • and task/role drift.

In the tested configurations, the results showed a clear trend in favour of Chmurka in the neutral critical-review condition. Chmurka more frequently challenged unsupported assumptions, false positives and internal inconsistencies in the source audits, while Gemini more frequently produced unsupported confirmations, false alarms or important omissions.

The professor-style instruction did not systematically improve performance. In several runs it reduced task compliance, and in two clear cases the model shifted from independently evaluating the audit to restructuring or rewriting it.

This is a small pilot study and does not establish universal superiority of one model over another. The sample size is limited, reasoning-effort settings were not fully controlled, model versions and configurations may differ, and some source audits contained excerpts rather than complete original conversations.

The results should therefore be interpreted as evidence of a trend in the tested configurations and as a basis for independent replication and further cross-model validation.

The next planned stage is blind cross-model validation, in which Gemini will evaluate Chmurka-generated reviews and Chmurka will evaluate Gemini-generated reviews using the same frozen evaluation protocol.

Tytuł:

Cross-Model Critical Audit Evaluation: Gemini vs Chmurka — A 12-Run Pilot Study

Keywords:

LLM evaluation, AI auditing, LLM-as-judge, cross-model evaluation, hallucination, epistemic calibration, task drift, instruction following, AI safety, Gemini

Files

Raport_koncowy_pilota_Gemini_vs_Chmurka_v1_0.pdf

Files (173.1 kB)

Name Size Download all
md5:f18ff2eaae0bca8542cc26a12bb1a375
173.1 kB Preview Download