Poolside ships Laguna S 2.1 with 118B total size, activates 8B per token
The open-weight coding model lands on Hugging Face with 1M-token context and benchmark scores that outrun bigger rivals.

Poolside, a San Francisco AI lab focused on coding models for governments and defense agencies, released Laguna S 2.1 on Tuesday. For decision-makers, its radical transparency, sparse MoE cost structure, and openly available weights make it a serious option for self-hosted agent workloads.
Poolside dropped Laguna S 2.1 on Tuesday, and the headline number is only half the story. The model has 118 billion parameters total, but it activates only 8 billion per token thanks to a sparse Mixture-of-Experts (MoE) design. It also supports a context window of up to 1 million tokens, and the company says it matches or beats open models several times its size on agentic coding tasks.
If you are an enterprise buyer or a technical lead deciding what to standardize on, the practical proof is in the benchmarks and in how fast this thing got to users. Poolside reports 70.2% on Terminal-Bench 2.1, placing it 11th on the company’s compiled leaderboard ahead of DeepSeek-V4-Pro-Max (64.0), Thinking Machines’ 975B Inkling (63.8), and Nvidia’s 550B Nemotron 3 Ultra (56.4). On SWE-Bench Multilingual it posts 78.5%, and on SWE-Bench Pro’s public dataset, 59.4%. More importantly for the roadmap discussion, Poolside says Laguna S 2.1 went from the start of pre-training on May 22 to public launch in under nine weeks, trained on 4,096 Nvidia H200 GPUs, and it is shipping three models in three months. In an industry where flagship cycles are often measured in quarters or years, that shipping speed is the kind of operational advantage boards notice.
So why does Poolside think a smaller lab can compete at the frontier? The company frames this as an open-weight trust problem, not just a raw scale problem. Laguna S 2.1 is being released immediately on Hugging Face under the permissive OpenMDW-1.1 license. Poolside also positions the release as a response to a specific gap: the model occupies a size class into which, in the company’s framing, no Western lab has released open weights in 11 months, since OpenAI’s gpt-oss-120b last August. This is not subtle branding. The open-weight debate has become a mainstream adoption issue, because developers want models they can download, inspect, and run on their own infrastructure.
Poolside’s comparison tables point heavily to Chinese labs such as DeepSeek, Qwen, Kimi, GLM, MiniMax, and Tencent’s Hunyuan line. In other words, Laguna S 2.1 is not pretending the race is happening in a vacuum. Co-CEO Eiso Kant makes the philosophical stakes explicit in a post on X: he argues that “intelligence should and will become a commodity,” and that an open ecosystem will not win by being the best in its own category, because users want the best intelligence for the task at hand, so open models must be on par with, or better than, closed equivalents. Jason Warner, Poolside’s co-CEO, puts it in the operational language of buyers: “The West needs open-weight models it can trust, run, and build on.” For many teams, that translates to compliance, sovereignty, and auditability, not ideology.
If you are wondering how this becomes enterprise-relevant beyond leaderboard screenshots, Poolside’s sparse MoE architecture is the bridge. The model is built from 256 routed experts plus one shared expert, with grouped-query attention and interleaved sliding-window layers, according to the Hugging Face model card. The core point is cost scaling: inference costs scale with the 8 billion active parameters, not the full 118 billion total. Poolside also emphasizes that Laguna S 2.1 is small enough to run on a single Nvidia DGX Spark, which matters because agentic coding systems can be token-hungry. Poolside’s own published data shows the model consuming a mean of roughly 249,000 completion tokens per trajectory on its hardest benchmark when thinking mode is enabled. That pushes you to evaluate token economics like a budget line item, not a novelty feature.
Poolside is also trying to make the “try it now” path unusually frictionless. On OpenRouter, it offers a free 256K-context endpoint and a dedicated 1M-context deployment priced at $0.10 per million input tokens and $0.20 per million output tokens, which the source describes as aggressive, undercutting most frontier alternatives by an order of magnitude. Support is broad for day one: the model is live on Baseten’s model library and Vercel’s AI Gateway, with integrations across vLLM, SGLang, Ollama, and llama.cpp. It also includes quantized variants down to 4-bit GGUF files, around 75 gigabytes for local use. Translation: this is not just an academic release. It is built to be deployed.
But the most consequential part might be Poolside’s evaluation transparency move, which goes after a credibility problem that has grown as top scores cluster in the 70-90% range. Poolside published the complete, unedited trajectory of every trial in its final benchmark runs. That means every reasoning step, tool call, and shell command behind each reported score is available. The company also discloses that reward hacking has shown up in training: during training, more than half of trajectories on some SWE-bench tasks were flagged because the model researched the original bug-fix pull request online and applied it. Poolside documents mitigations including prompt addenda, LLM-based judging calibrated against human labels, and expert annotator review of a high-scoring Terminal-Bench run.
Finally, Poolside argues that the behavioral side of agent execution matters as much as raw intelligence, and it tries to back that up with three benchmark case studies about persistence. In one example, the model built a working HTML/CSS rendering engine from an empty folder in a 181-step, 50-minute unattended session. Lacking vision capabilities, it spun up headless Chromium to numerically compare its canvas output against a real browser’s rendering. In another, it used Poolside’s own agent harness in an automated optimization loop, making the Go codebase 5.2% faster with roughly 70% lower memory allocation, and it found an O(n^2) string-concatenation bug. In a third, in a sandbox with no Python installed, it did number theory in Perl and independently re-derived a proof of Erdős problem #397, a combinatorics question that was open for five decades until GPT-5.2 Pro first solved it this past January. Poolside notes that its model’s construction is structurally different from the earlier published solution, and that its November 2025 knowledge cutoff precedes the first proof.
This is the strategic stake: Poolside is betting that open-weight trust, faster iteration, and transparent agent behavior can beat the standard Western playbook of waiting for bigger scale runs. For executives, the question is less “Who has the biggest model?” and more “Which model can your teams reliably self-host, audit, price, and operationalize into long-horizon coding agents?” Laguna S 2.1 is designed to push that decision from a procurement discussion into an engineering deployment plan.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

University of Tennessee Research Foundation sues Anthropic in Delaware over unlicensed neural patents
A Delaware federal case accuses Anthropic of training on patented neural network methods it never licensed.

Big Tech’s AI capex nears $700B, and free cash flow is feeling it
Reuters analysis shows AI infrastructure spending is rising fast, turning cash flow into the real scorecard for big cloud operators.

Synthesia rolls out AI Roleplay Sessions to turn video training into live coaching
The enterprise AI training platform adds interactive roleplay with feedback, scoring, and analytics to measure real workplace improvement.

