Atlas-Finance: Evaluating AI Agents Inside a Bank

(joinhandshake.com)

2 points | by cjbarber 5 hours ago

1 comments

  • cjbarber 5 hours ago
    I found this interesting.

    > We evaluate 11 frontier models (see Figure 1), all run within the OpenCode agentic harness. Claude Opus 5 performs the best, yet still only manages to pass 12.3% of tasks. Claude Fable 5.1 and GPT-6 Astra are close behind, but the other eight models have significantly lower pass rates.