Published March 3, 2026 | Version v1

Scientific Coding with AI - SciCode Bench Insights & Agentic Workflows

Authors/Creators

Description

Talk given to the CASS user/developer experience working group by Andrew Schmeder of Lawrence Berkeley National Lab on March 3, 2026. Recording available here

Can LLMs actually perform “PhD-level” tasks - specifically in scientific coding - as claimed by AI companies? Recent advances have enabled the majority of UI and infrastructure code to be automated using AI, but can it write scientific code? In this short talk, we will review the results from running the SciCode benchmark on 60 different model configurations over the past 9 months on Berkeley Lab’s CBorg AI inference gateway. Insights regarding evals, optimizing inference costs, performance of open-weight versus commercial flagship models, and measuring the rate of model improvement will be discussed. In the second half, we will look at a demo of an autonomous scientific agent that can perform data discovery, data transfers and data analysis via a chat interface.

Files

CASS - SciCode 2026-03-03.pdf

Files (2.0 MB)

Name Size Download all
md5:dc9edbcbfaf01254a0bc42cf22ef5518
2.0 MB Preview Download