Skip to content
yceffort
PostsSeriesTagsAbout🧪 Research
KO

Tweaks

theme
accent palette
film grain
minimal mode
◆ SERIES · 1 POSTS

Deciding Whether to Swap the Model Behind a Coding Agent

Examining how SWE-bench builds and grades its tasks, designing an experiment to compare models in Claude Code, and deciding whether to switch based on resolved rate and cost.

README

A record of finding out whether the model used in Claude Code can be replaced with an open-weight model. Part 1 analyzes how SWE-bench tasks are built and graded and sets out the principles for comparison. Part 2 will set up the actual execution environment and experimental conditions, and part 3 will decide whether to switch models based on resolved rate, cost, and failure cases.

01 ITEMS

All posts

in order
01

How to Compare the Models Behind a Coding Agent with SWE-bench

#ai
How do you measure the model performance of a coding agent?
2026-10-1028 min · read
mailMail icongithubtwitter
yceffort
•
© 2026
•
https://yceffort.kr