On July 1, 2026, Together AI closed a new funding round at an $11 billion valuation, co-led by General Catalyst and Prosperity7, with participation from Saudi Arabia's Public Investment Fund (PIF), NVIDIA, Salesforce Ventures, and others. Less than five months earlier, the company was valued at $3.3 billion — meaning it more than tripled in about 150 days, placing it directly in the top tier of US AI infrastructure unicorns.
What Together AI does can be summarized in one sentence: it is not a model company — it is a wholesaler of AI inference compute. It owns and operates its own GPU clusters, and customers pay per token for compute. The underlying model can be any open-source model — Llama 4, DeepSeek, Qwen 3.x. Running a model on Together AI is logically the same as renting a generator at a factory to produce your own power.
Why an $11B valuation?
Because Together AI is the biggest beneficiary of the open-sourcing of AI inference. OpenAI, Anthropic, and Google lock top models behind closed APIs, but developers, enterprises, and research institutions worldwide need open-source models — especially in markets like China, Europe, and India that are unwilling or unable to use closed models. Together AI packages top open models like Llama 4, DeepSeek-V3, and Qwen 3 into token-billed inference APIs, so customers never have to deal with GPU procurement, model deployment, or scaling.
Three layers of support for the $11B valuation
Layer one: the 'neutral infrastructure' position. With OpenAI, Anthropic, and Google locking in customers, there is clear demand for a 'neutral inference layer' — customers want models they can swap anytime and compute they can migrate anytime. Together AI is currently the largest neutral inference provider.
Layer two: the customer list. Publicly disclosed customers include Salesforce, Zoom, Cartesia, and Hedra — not large enterprises that buy 1,000 GPUs themselves, but mid-sized companies that consume AI on demand. This customer profile overlaps heavily with Microsoft's July 2 'Frontier Company' customers, but Together AI bills per token while Frontier Company bills per project, so the two do not directly conflict.
Layer three: open-source ecosystem binding. Together AI is one of the officially recommended inference platforms in English-speaking markets for top open models like Llama, DeepSeek, Qwen, and Mistral. Whenever an open model ships a new version, Together AI deploys it within 24–48 hours, with zero switching cost for customers.
What does this have to do with AI coding?
A great deal. The core bottleneck for AI coding tools is no longer 'which model to use' but inference cost. The margins of tools like Cursor, Claude Code, and Copilot depend directly on their per-thousand-token inference cost. Cursor's core 2026 strategy is 'multi-model routing' — automatically switching between Claude Sonnet 5, Grok 4.5, and Llama 4 based on task difficulty, using expensive models for hard tasks and cheap models for easy ones.
The foundation of multi-model routing is exactly a neutral inference platform like Together AI. If Cursor ran everything on OpenAI's API, it would be locked into OpenAI's pricing; by routing 30% of its inference traffic to Together AI's Llama 4, it can cut overall costs by 40–60%. This is a fundamental restructuring of AI coding tools' cost structure.
Going further: Anthropic's July 1 release of Claude Projects / Claude Code enhancements is built on Anthropic's in-house models, but Anthropic has also open-sourced the Claude Code SDK — developers can combine the Claude Code SDK with Together AI's Llama 4 inference API to build their own 'Claude Code-like' tools at one-third to one-fifth the cost of the original.
This is an early signal of vertical stratification in the AI coding tool market:
- Top tier (flagship products): Anthropic Claude Code, Cursor, GitHub Copilot Workspace — strongest closed models, $20–200/month per seat.
- Middle tier (customizable products): internal tools built on Together AI + open-source models — variable pricing, but controllable cost.
- Bottom tier (free tools): VS Code + Continue.dev + Ollama local models — free but limited capability.
The risks, honestly stated
1. Position under attack. Together AI's core edge is 'neutral + open-source,' but that is exactly the position OpenAI, Anthropic, and Google most want to eliminate. All three have their own inference services and are willing to burn money at negative margins. If OpenAI cuts GPT-5.x inference API prices to 70% of Together AI's in late 2026, Together AI's customer base could loosen immediately.
2. Token-billing economics under GPU deflation. Per-token billing is healthy as long as GPU prices are stable, but H100/H200 prices entered a rapid decline in 2026. If Together AI signed long-term supply contracts (many large customers do), its own cost decline may not keep pace with market price declines — a classic supply-demand timing mismatch.
3. Geopolitical tension. The dual entry of Prosperity7 (Aramco Ventures) and PIF means Together AI has secured a ticket to Middle Eastern 'sovereign AI' capital — but Middle Eastern capital has its own preferences on return horizons and compliance. Over the next 12 months, Together AI may face tension between US customers and Middle Eastern investors.
Overall, Together AI's $11B valuation is one of the most significant marker events of 2026 in the AI infrastructure space. It signals that 'AI inference' has evolved from a subordinate of models into an independent track. AI coding, AI agents, embodied AI — every application-layer innovation that requires massive inference compute will grow on top of it.
Source: Together AI official blog, 'Together AI raises at $11B valuation led by General Catalyst and Prosperity7,' 2026-07-01, https://www.together.ai/blog/together-ai-raises-11b-valuation