Build a small personal eval harness on your own repo, because the public leaderboard is answering a different question than the one you're asking.
Part of: Coding on Open Models