How models
actually code.

Public scores for models available in Boongle IDE — first-party, catalog, and API tiers. We combine published industry benchmarks, independent evals, and BoongleBench runs in our agent harness.

15models tracked

Auto, Boongle 1.8 Flash, Grok 4.6, Claude, Gemini, GLM, GPT, Mimo, DeepSeek V4.1 Flash, Gemini 3.8 Flash, and more.

5benchmark axes

SWE-Bench Multilingual, Terminal-Bench 2.1, agent coding evals, speed, and cost efficiency.

Boongle 1.8 Flash

Our latest first-party model for agentic coding in Boongle IDE — tuned for multi-file edits, fast tool loops, and everyday refactors inside the editor. Read the release

  • 78.6 SWE-Bench Multilingual
  • 71.5 Terminal-Bench 2.1
  • 61.9 Agent bench v3.1
78.4BoongleBench composite

Strong all-round scores with high speed and cost efficiency on Boongle IDE plans. Pick it when you want flagship quality without switching models.

What we measure

SWE-Bench Multilingual

Real bug fixes across multilingual open-source repos. Higher is better. Source: public SWE-Bench Multilingual leaderboard and vendor reports.

Terminal-Bench 2.1

Agentic terminal workflows — install deps, run tests, debug CLI output. DeepSeek V4.1 Flash and Gemini 3.8 Flash lead here.

Agent coding

Public agent-coding evals where available (e.g. agent bench v3.1, DeepSWE v1.1). The exact benchmark is labeled on each model row.

BoongleBench composite

Weighted blend of available axes (coding 45%, terminal 25%, speed 15%, cost 15%). Missing axes are skipped — never guessed.

Every model.
One table.

# Model Composite SWE Multi TB 2.1 Agent Speed Cost

Scores combine publicly reported vendor benchmarks, independent evals (e.g. Vals AI, Artificial Analysis), and BoongleBench internal runs in the Boongle IDE agent harness (March 2026). Third-party benchmark names belong to their respective publishers. Different models are often evaluated with different harnesses — compare within the same column, not across unrelated evals. We update this page when new catalog models ship.