Open-weight model dashboard

Track model quality, licensing, local deployment fit, and evidence confidence across the coding-model landscape.

Updated Jul 27, 2026
Tracked models
0
Active comparison set
American-controlled
0
Current research focus
Local-ready
0
Ollama / vLLM / GGUF paths
Immediate test candidates
0
Needs harness validation

Agentic coding benchmark view

Reported scores emphasize coding-agent strength, but evidence quality varies by benchmark harness and verification depth.

Evidence posture

Verified factArchitecture, parameter counts, context, quantized variants, and runtime claims with primary source backing.
Company claimHeadline benchmark superiority, DFlash acceleration, and broad comparative leadership language.
UnresolvedTraining data provenance, clean apples-to-apples benchmark replication, and long-term openness durability.

Decision notes

Laguna XS 2.1Strong local coding candidate. Benchmark conditions require independent validation before fleet standardization.
Qwen-class competitorsRemain reference points for coding performance and cost-efficiency in the global field.
Dashboard modelStore evidence confidence separately from raw scores to avoid leaderboard over-trust.

Model registry

Single-table intake for model, developer, geography, architecture, hardware fit, benchmark posture, and recommendation state.

Model Control Architecture Total / active Context License SWE-V Terminal Local support Verification Recommendation

Laguna XS 2.1 snapshot

DeveloperPoolside
Control / developmentFrance / United States
Release date2026-07-01
Total / active33B / 3B
ArchitectureMoE, 256 experts + shared expert
Context262,144 tokens
LicenseOpenMDW-1.1
Local runtimesOllama, vLLM, SGLang, llama.cpp
Current recommendationTest immediately

Key risks

Benchmark independenceLimited
Training provenancePartially disclosed
Runtime maturityGood, still stabilizing
General-purpose transferNot yet established
Commercial openness durabilityUnknown
Operational upsideHigh if local coding holds

First test plan

Checkpoint 1Official GGUF / Ollama
Checkpoint 2FP8 or NVFP4 GPU run
Harness test 1Repo bug fix
Harness test 2Tool-call reliability
Harness test 3Structured output
Harness test 4Long-context refactor
Harness test 5Strict terminal step cap

Dashboard data model

Per-model recordOne row per checkpoint
Evidence fieldClaim / fact / verified
Hardware tierMinimum vs practical
Benchmark metadataHarness, attempts, steps
Recommendation stateFleet intake ready
Last updateDate-bound research refresh

How this dashboard should be used

Do not collapse research into one leaderboard rank. Track model strength, deployment cost, license friction, and verification quality as separate dimensions.

PrimaryModel recordArchitecture, context, checkpoints, runtimes, licensing label, and release date.
ClaimVendor benchmarkStore harness, attempts, context, tools, and patching notes beside every score.
VerifiedIndependent testPromote recommendation state only after internal or external matched-condition reproduction.

Field groups

IdentityModel, developer, region of control, development footprint, release date.
TechnicalTotal and active parameters, architecture, attention design, context, quantizations.
OperationalMinimum RAM / VRAM, practical hardware tiers, tokens per second, local runtime support.
EvidenceBenchmark source, harness details, independent verification, recommendation state, last update.