There is a newer version of the record available.

Published May 9, 2026 | Version 0.1-bilingual

macbench: A macOS-Native Computer-Use Benchmark for Autonomous Agents

Authors/Creators

Description

The first publicly published macOS-native computer-use benchmark. 369 task slots across 15 categories (Finder, Safari, Mail, Notes, Calendar, Reminders, Settings, Terminal, Pages, Numbers, Keynote, Music, Photos, Maps, Multi-app), agent-agnostic Go runner, dual scoring (IMPLEMENTED + STRICT), per-task PID-snapshot isolation. First reference run: kinclaw v1.15.0 + Kimi-K2.5 = 67.3% IMPLEMENTED. Documents the full 49.3 -> 62 -> 67.3 debugging trajectory as methodology contribution.

Note (2026-05-09): This version bundles English + 中文 in a single PDF (English first, then Chinese), generated directly from the canonical Markdown source files.

Files

macbench.pdf

Files (1.4 MB)

Name Size Download all
md5:3c375551d047bc553b8c399899cf701e
1.4 MB Preview Download

Additional details

Related works