Published May 9, 2026
| Version 0.1-bilingual
Preprint
Open
macbench: A macOS-Native Computer-Use Benchmark for Autonomous Agents
Authors/Creators
Description
The first publicly published macOS-native computer-use benchmark. 369 task slots across 15 categories (Finder, Safari, Mail, Notes, Calendar, Reminders, Settings, Terminal, Pages, Numbers, Keynote, Music, Photos, Maps, Multi-app), agent-agnostic Go runner, dual scoring (IMPLEMENTED + STRICT), per-task PID-snapshot isolation. First reference run: kinclaw v1.15.0 + Kimi-K2.5 = 67.3% IMPLEMENTED. Documents the full 49.3 -> 62 -> 67.3 debugging trajectory as methodology contribution.
Note (2026-05-09): This version bundles English + 中文 in a single PDF (English first, then Chinese), generated directly from the canonical Markdown source files.
Note (2026-05-09): This version bundles English + 中文 in a single PDF (English first, then Chinese), generated directly from the canonical Markdown source files.
Files
macbench.pdf
Files
(1.4 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:3c375551d047bc553b8c399899cf701e
|
1.4 MB | Preview Download |
Additional details
Related works
- Is supplemented by
- Software: https://github.com/LocalKinAI/macbench (URL)
- Software: https://github.com/LocalKinAI/kinclaw (URL)