All notes

#198 Open source AI in 2026: The founder briefing

July 25, 2026·19 min read

#198 — Open source AI in 2026: The founder briefing

Open source AI just crossed into serious business territory, and the biggest opportunities now sit above the model, not inside it.

Open source AI in four numbers

  • Open models have closed the gap to the top closed models to 3.3 points, down from 8.04 points in January 2024, and briefly hit near zero in February 2025 before reasoning models pulled ahead again.
  • Inference pricing for GPT-4-class performance dropped 50x in three years, from $20 to $0.40 per million tokens, a steeper fall than the historical PC-compute or dotcom-bandwidth curves managed over the same stretch.
  • Open weights now handle about a third of all tokens on OpenRouter, a major model-routing marketplace, and Mistral's annual revenue jumped from $20M to $400M in a year.
  • This is a multi-hundred-billion-dollar commercial layer now, and as model pricing heads toward zero, the competition is moving up to the software wrapped around the model.

Your 2023 impression of open models is out of date

Llama 2 had genuine quality problems, hallucinations, and fine-tuning failures that sent teams back to GPT-4, but that's now two model generations behind us. By January 2025, DeepSeek-R1 matched OpenAI's o1-1217 on math and reasoning benchmarks (97.3% vs about 97% on MATH-500, 79.8% vs about 79% on AIME) at roughly 27 times lower cost ($0.55/$2.19 vs $15/$60 per million tokens), released under an MIT license. By April 2026, DeepSeek-V4 Pro scaled to 1.6T parameters and a 1M-token context window for $0.435 per million input tokens. At this point the gap left is the software layer, not the model.

That 3.3-point average hides real variation across tasks:

  • Near parity: coding, instruction-following, general knowledge
  • Closed models still ahead: reasoning, long-context retrieval, agentic tasks

Open wins usage. Closed wins the money

The five highest-volume models on OpenRouter are all open weight:

  • DeepSeek V4 Flash (18.4T tokens/month),
  • Xiaomi's MiMo-V2.5 (14.9T),
  • Tencent's Hy3 (14.8T),
  • MiniMax M3 (14.3T), and
  • a stealth release later identified as Meituan's LongCat-2.0 (11T).

Among the top nine models by monthly tokens, Chinese-built open models moved about 82T against roughly 23T for US-built closed models, more than a 3.5-to-1 split. Counted by request volume instead of tokens, closed US providers still lead, with Google alone taking 28.7% and Anthropic plus OpenAI around 20% combined. Open's strength concentrates in coding and agentic workloads, where token counts run highest.

Here's the number that stands out most: open models handle roughly 80% of usage on OpenRouter but bring in only about 4% of the revenue.

The single most-used model by volume, DeepSeek V4 Flash, ranks sixteenth by measured intelligence, and Anthropic holds three of the four most-capable models while routing only about an eighth of weekly tokens, so usage clusters at the cheap commoditizing layer while revenue collects higher up. Closed models charge about six times more per call for roughly 90% capability parity, and the Linux Foundation estimates $24.8B in savings companies are leaving on the table by not switching.

About 30% of developers who choose open cite lower cost and privacy as their top reasons. Developers rarely pick a single lane:

  • 50% run both open and closed models,
  • 29% run open only,
  • 21% closed only, and
  • open users average more use cases per developer, 5.1 versus 4.6, with 78% pairing specialized, task-tuned open models alongside general-purpose ones.

Capability isn't what's holding open models back

89% of firms use open components and 79% of developers use open models, yet only 51% of open models reach production, compared to 63% for closed.

Here's the number founders should sit with: internal builds hit production only 33% of the time versus 67% for vendor-partnered deployments, and enterprise pilots showing measurable financial impact land at just 5%.

Scale doesn't close this gap either. Open production rates barely shift by company size (53% small, 55% mid-size, 57% enterprise), while closed climbs sharply with resources (54%, 66%, 73%). Money fixes closed deployment problems. Open deployment problems are still waiting on the ecosystem to build the missing pieces.

What blocks adoption and pushes people to churn, ranked by how often developers cite it:

  • infrastructure and compute costs (27%),
  • security and compliance concerns (26%),
  • ongoing maintenance burden (up 11 points among people who churned),
  • integration difficulty (up 11 points), and
  • thin documentation (up 9 points).

These patterns repeat across regions, with security concerns peaking in South Asia at 39%. Across a stack-maturity assessment scoring 48 subcomponents on nine criteria, every layer scores strong on community activity and raw capability but weak on standardization and enterprise readiness. Model code and weights score highest overall (around 3.7 to 4.0 out of 5), while safeguards score lowest (2.64), and that standardization-and-readiness weakness repeats down every single layer.

Who's actually turning this into revenue

CompanyMetricDetail
Databricks$5.4B revenue run-rateOver 65% YoY growth, weighing a raise at a $165B+ valuation
Mistral$400M ARRUp 20x in a year, in talks for 3B at a 20B valuation
DeepSeek$220M ARRRaised $7.4B at a $50B+ valuation

Five revenue models are proven at scale:

  • hosted inference,
  • enterprise platforms,
  • on-prem licensing,
  • fine-tuning services, and
  • harness tooling.

On the funding side, DeepSeek has raised $7.4B total, Moonshot AI $3.9B, Mistral $3.05B, Reflection AI $2.13B, Cerebras $2.1B, and Cohere $1.7B, spread across models, inference, tooling, and compute.

Every major hyperscaler holds a stake here:

  • Microsoft and Amazon invested in Mistral and Hugging Face respectively,
  • NVIDIA backs nearly the whole inference stack (Mistral, Together, Cohere, Fireworks, Baseten, Hugging Face, Replicate) while also shipping its own Nemotron models,
  • Google ships Gemma while backing Hugging Face,
  • IBM ships Granite,
  • Meta ships Llama.

The largest disclosed strategic rounds include ASML's $1.4B into Mistral, Tencent's $1.4B into DeepSeek, and CATL's $700M into DeepSeek. Consolidation is already underway: CoreWeave bought Weights & Biases for $1.7B, Databricks bought MosaicML for $1.3B, NVIDIA bought Run:ai for $700M, AMD bought SiloAI for $665M, and Cohere's deal for Aleph Alpha landed near $20B in enterprise value.

One gap founders should notice: across seven stack layers, private capital pours heavily into foundation models, inference, and compute, but data and datasets, and safety, evaluation, and governance, the trust layer, run almost entirely on philanthropy and government grants with essentially zero private VC or strategic corporate money.

Metered pricing keeps blowing up budgets

Three enterprise stories make the case.

  • Microsoft is cancelling most of its Claude Code licenses by June 30, 2026, after token billing consumed a division's annual AI budget in months.
  • Uber exhausted its entire 2026 AI coding budget in four months, with engineers billing $500 to $2,000 a month before the company capped spend at $1,500 per tool per employee.
  • Stripe cut inference costs 73% by self-hosting open models on vLLM, running 50 million daily API calls on a third of its GPU fleet, a fixed cost it controls and can plan around.

By June, Microsoft, the world's largest software company, was itself exploring Azure-hosted DeepSeek V4 for its heaviest Copilot workload, working around its own partner's billing meter. In both failures, the trigger was the same: a shift to usage-based pricing that put the vendor, not the customer, in control of unit economics.

This tracks the same pattern as cloud repatriation. Moving a petabyte out of AWS S3 costs $90K to $120K in egress fees, roughly 80% of enterprises are now pulling some workloads back on-prem, GEICO's cloud costs ran 2.5 times over expectations before it repatriated, and 37signals projects over $10M in savings over five years after leaving the cloud entirely. Open weights are portable because you control the exit; self-hosted infrastructure is a fixed cost with no forced lock-in. Closed model APIs recreate the same trap, where the vendor controls pricing and there's no clean migration path.

Sovereignty went from theory to a live risk

In June 2026, a real sequence played out in public. Anthropic shipped Fable 5 and Mythos 5 on June 9. Three days later, on June 12, a US export-control order cut off access effective immediately, barring any foreign national from using either model, inside or outside the US, including Anthropic's own foreign-national staff. Since nationality couldn't be verified in real time, selective compliance was impossible, so both models went fully dark for everyone. Partial clearance came June 26, restoring Mythos to roughly 100 vetted US critical-infrastructure organizations, and full controls lifted June 30, with Fable 5 restored globally July 1, nineteen days after access was cut. A model you rent can be switchedoff by someone else. A copy running on hardware you own cannot.

The same logic plays out at national scale. China's open-weight strategy is explicit industrial policy under its AI-Plus directive and current Five-Year Plan: release open weights, offload inference onto users' own hardware, and let global demand, especially in the Global South, diversify away from the US stack as a hedge against semiconductor export controls.

  • Alibaba's Qwen has racked up 942 million cumulative Hugging Face downloads against Meta's Llama at 476 million, and in February 2026 Qwen alone out-downloaded the next eight organizations combined.
  • Chinese open models went from under 2% to over 45% of weekly OpenRouter tokens between late 2024 and April 2026, now 61% of traffic among the platform's ten most-used models, with DeepSeek alone counting 26,000-plus enterprise accounts and appearing in 58% of new 2025 AI-startup stacks.

The caveats are real, training-data opacity and hard-coded refusals among them, and at least eight jurisdictions ban the hosted apps outright, but enterprises keep adopting the underlying weights anyway, self-hosted or through Western endpoints.

Europe is treating open weights the same way, escalating from France's 109B AI infrastructure commitment in February 2025, through AI Act exemptions for qualified open source systems in August 2025, to the EU's EUROPA award in June 2026 for a sovereign 400 billion parameter open model spanning all 24 EU languages, and Portugal and Germany shipping their own open systems in July 2026. Canada committed $890M to a sovereign public AI supercomputer as part of Prime Minister Carney's "AI for All" initiative in June 2026, and Cohere's Nick Frosst has framed his company's decision to open-source the 218-billion-parameter Command A+ explicitly in sovereignty terms. India added over five million new GitHub developers in 2025 alone, reaching 21.9 million total, the fastest-growing developer base in the world, while subsidizing 38,231 GPUs at roughly 40% below market rate and accounting for 13.6% of DeepSeek's monthly active users. Saudi Arabia has committed $77B and South Korea $71.5B to national AI infrastructure, and over 80 jurisdictions now have AI policies on the books, with 47 restricting foreign processing for critical workloads.

Where the real competition is happening now

Picture the agentic harness, orchestration, tools, memory, sandboxes, and permissions, as a step up from the browser: code on your side that negotiates with the world on your behalf, the way the browser once negotiated with servers. This is already a real product category. LangChain holds over 126,000 GitHub stars and 60% developer share in orchestration, and MCP hit 97 million monthly downloads with over 10,000 active servers in its first year. The market is also starting to consolidate upward into "metaharness" layers, like Databricks' open-sourced Omnigent, that wrap multiple separate agent ecosystems under one plane of governance.

Breaking the harness down by what it actually does, from the bottom up:

Sub-layerStateKey players
Orchestration loopOpen-ledLangGraph, CrewAI, AutoGen stateofopensource
MemoryOpen-ledMem0 (47,000+ stars), Letta, Zep stateofopensource
Tools/interopOpen standardMCP, A2A stateofopensource
Sandboxes/evalMixedE2B, Daytona, Modal, Langfuse stateofopensource
Write permissionUnsolvedNo portable standard exists stateofopensource
GovernanceEmergingOmnigent, OPA, agent governance toolkits stateofopensource
  • The model itself: open or closed, swappable, and commoditizing toward zero
  • Control (drives the loop): the orchestration loop, the reason-and-act cycle that turns a model into an agent, led by open tools like LangGraph, CrewAI, AutoGen, and LlamaIndex
  • Reach (connects and remembers): tools and context through MCP, agent-to-agent communication through A2A, and memory through Mem0, Letta, and Zep
  • Action (does things safely): sandboxes and execution (E2B, Daytona, Modal), permission and identity (the write surface, still the unsolved gap), and eval and observability (Langfuse, Phoenix)
  • Surface (meets the user and money): interface standards like AG-UI and A2UI, and payment and metering protocols like x402, AP2, and UCP
  • Govern (one plane over many harnesses): stateful policy tracking what a session already did, a registry and lineage of which agent did what, and budget controls with kill switches, through tools like Omnigent, OPA, and emerging agent governance toolkits

Every sub-layer already has real products in it. Orchestration and memory are open-led, interop runs on open standards, sandboxes and evaluation are mixed, and permission is the youngest and least standardized category, the clearest open gap in the whole stack.

The trend to watch closest: model makers are pulling this software layer in-house, and independent tools are losing ground fast.

  • On Terminal-Bench 2.0 in May 2026, a third-party scaffold running Anthropic's own weights beat Anthropic's official Claude Code tool by 21.8 points on the identical model (79.8% vs 58.0%), meaning the harness was outperforming the weights themselves.
  • Eight weeks later, on Terminal-Bench 2.1, that gap had shrunk to about 3 points: Codex CLI on GPT-5.5 scored 83.4%, Claude Code on Claude Fable 5 scored 83.1%, and the best independent harness on Fable weights only reached 80.4%.

On every model where both versions exist, the lab's own tool now wins, and no open model appears in the verified top tier. A tool built tightly around one lab's model performs worse on everyone else's, and that's becoming a form of lock-in as a side effect of optimization, not a deliberate plan, forming a real moat. On a neutral scaffold, though, price differences stay wide: GLM 5.2, an open model, scores 67.79% at $0.43 per task, within about a point of Claude Opus 4.7's 68.54% at $1.98, roughly a fifth of the cost. Integration also buys the labs something else: usage exhaust that trains whoever owns the harness.

Adoption of the shared standards here is running well ahead of governance. MCP grew from about 2 million monthly downloads at its November 2024 launch to 97 million by early 2026, with OpenAI adopting it in March 2025, OAuth 2.1 built into the spec by June 2025, 28% of the Fortune 500 running it in production by March 2026, and Anthropic donating it to the Linux Foundation's new Agentic AI Foundation in December 2025, alongside Block's goose and OpenAI's AGENTS.md, with platinum members including AWS, Google, Microsoft, and OpenAI. But only about 21% of companies report mature agent governance, and over 30 CVEs were filed against MCP in just the first eight weeks of 2026, right after it became Linux Foundation infrastructure. MCP hardened onto OAuth 2.1 and A2A standardized signed Agent Cards, but both stop at authentication, proving who's asking, without solving authorization, deciding what they're allowed to do.

Recent incidents prove closed platforms aren't automatically safer. Anthropic's own Slack MCP integration, Microsoft Copilot ("EchoLeak"), Salesforce Agentforce ("ForcedLeak"), and ServiceNow ("BodySnatcher") all suffered critical zero-click exfiltration bugs between June and October 2025, and all four shared the same root cause: retrieval was checked, but output never was.

Interestingly, 41% of developers associate closed models with privacy and security versus 29% for open, but that gap tracks who carries the operational burden, since closed APIs ship safeguards on by default while open deployments require you to wire in the same controls yourself. It measures who does the work, not where the actual risk sits.

Closed models still hold a real, measurable lead in a few places:

  • a 14-point gap on Terminal-Bench agentic-terminal tasks (GPT-5.5 at 82.0% vs DeepSeek-V4-Pro-Max at 67.9%),
  • long-context accuracy (Gemini 3 hits 89% multi-needle recall at 1M tokens versus GPT-5.5 at 74%, Claude Opus 4.7 at 56%, and DeepSeek V4-Pro at just 41%), and
  • packaged compliance like SOC 2, HIPAA, and zero-data-retention agreements that closed providers bundle by default.

DeepSeek scores an F, 0.37 out of 4, on the FLI AI Safety Index, and there's also a basic accountability difference: pay a closed vendor and someone else carries liability when things break, while self-hosting open weights makes it entirely your own.

Two assets worth building for yourself

Here's the core architectural idea: model value commoditizes with time and prices toward zero, while memory value compounds, growing with every interaction in production. A rented model can be deprecated by whoever built it. Memory held on your side of the firewall can't be, and it can't be reacquired just by switching vendors.

  • portable formats, like plain markdown under version control synced to native contacts and spreadsheets that no provider governs;
  • retrieval that checks your own private memory before the open web, so the system isn't filling private gaps with confident public guesses; and
  • append-only storage that never overwrites what a user actually entered.

The other genuinely open problem is permission, and it splits cleanly into two surfaces with very different stakes.

  • Reads are reversible and low-consequence, so they can largely be permitted by default.
  • Writes, sending a message, spending a budget, modifying a record, executing a transaction, carry costly or irreversible side effects, and that's where confirmation, thresholds, cost caps, and revocation all need to concentrate.

Zero portable write-permission standards exist across the ecosystem's 12 frameworks, 10 harnesses, and 3 peer protocols. MCP hardened onto OAuth 2.1 and A2A standardized signed Agent Cards, but authentication isn't authorization, and both stop short of it. A handful of players are trying to bridge the gap on identity and auth (Okta, WorkOS, Auth0, Stytch, Arcade) and on policy engines (OpenFGA, Cedar), and CoSAI, an industry security coalition, ranks consent fatigue as a top-tier threat: users approve most prompts, and the prompts that matter most are exactly the ones authorizing action. The emerging meta-harness architectures that enforce stateful policy above any single agent, gating the next write based on what a session already did, look like the most likely place for a durable permission model to actually form.

Five bets worth making now

Each bet here is scored on the same three things: what the current data shows, how much time is left on the clock, and what standing still costs you.

  • Build the open harness, co-designed with open weights and tuned to them the way Codex is tuned to GPT-5.5, either general-purpose or built for a vertical the frontier labs aren't targeting. Frontier-model funding rounds absorbed hundreds of billions in 2026 while the open-harness category took a rounding error, and this window closes once the closed stacks weld model and scaffold into one rented product
  • Own the memory layer, portable and append-only, behind your own firewall. Open memory tooling already runs at scale, with Mem0 past 47,000 GitHub stars alongside Letta, Zep, and LangMem, and every quarter spent on a closed endpoint hands your one appreciating asset to a vendor you don't control
  • Solve portable write-permission, the unsolved hole at the center of the harness. MCP already crossed 10,000 servers and 97 million monthly downloads while governance maturity sits at only 21%, meaning the plumbing scaled but the lock never followed, and whoever standardizes it first decides whether renting stays the "safe" default for everyone else
  • Break the meter by second-sourcing your model now, while pricing is still cheap and boring. Self-hosting breaks even above roughly 8,000 conversations a day, and Uber's own ride fares rose about 92% once riders had reorganized their lives around the subsidized price. Introductory model pricing is expected to end around 2027 to 2028, once providers have gone public and the discounts run out
  • Make the open default plural by funding alternatives before one origin becomes the only supplier. Qwen's 942 million downloads already dwarf Llama's 476 million, and Chinese models make up over 45% of weekly tokens. Open weights still hide their training data, alignment choices, and refusal patterns, and when one origin supplies the default, everyone downstream inherits that single builder's blind spots with nothing to check them against

What would flip this analysis

Four lanes are worth tracking between now and the next assessment, and any two moving the wrong direction together would be the signal to revisit this whole picture.

  • On capability and adoption, watch the 3.3-point gap and open's OpenRouter token share, especially in agentic coding; this reverses if token share stalls while the reasoning gap widens.
  • On the harness, watch the lab-versus-independent Terminal-Bench spread and whether a portable permission spec ever appears under the Agentic AI Foundation's governance; this reverses if the lab-harness lead widens again or a closed platform sets the permission standard first.
  • On market structure, watch open-lab economics, ARR, funding rounds, and IPOs like Zhipu and MiniMax's, against the 2027-28 metered-pricing breakpoint, with sovereign government funding as a counterweight; this reverses if sovereign funding lapses or open-lab economics fail to scale.
  • On trust and safety, watch misuse capability and how easily safety tuning strips away, particularly around hard-friction harms, against whether the US NTIA's "monitor, don't restrict" posture holds; this reverses with a major misuse event or a shift toward restriction.

Open source AI is necessary here, but not sufficient on its own. It still needs hard friction in the places where harm concentrates, and owning the engine settles who captures the gains, not who owns the underlying data. What it does is put the engine within reach and make the fight winnable for people outside the largest labs.

more than just words|

We're here to help you grow better at every stage of the climb.

let's go to market

Whether you're finding problem-market fit, refining your positioning, shipping product, or scaling go-to-market we're built for every stage of the journey.