Baseten — Private AI Inference at $13B: Growth Priced, Margin Unknown
Published · entry price $ · machine-generated by the QuantLogix Thesis Engine and graded publicly at T+7/30/90 days · BASETEN charts & signals →
Baseten's $600M annualized run-rate (Mar 2026) and ~20x YoY growth make it the fastest-scaling private inference platform, but the $13B Series F valuation embeds 21.7x revenue at a margin profile that is almost certainly sub-50% given rented GPU COGS. The thesis turns entirely on whether dedicated inference sustains pricing power as GPU scarcity eases — a question no available data resolves. WATCH pending gross margin disclosure or a credible margin expansion signal.
Thesis
- Baseten's annualized revenue run-rate reached ~$600M by March 2026, up from ~$200M ARR in December 2025
- Baseten's Series F valuation of $13B implies ~21.7x annualized revenue, a premium that requires sustained 70%+ gross margins to justify against SaaS comps
- Inference peer gross margins run ~50% per Sacra data on Fireworks, well below the 70%+ SaaS benchmark, because rented GPU capacity sits in COGS
- Baseten surpassed its full-year CY26 revenue forecast by end of Q1 2026, implying forward revenue may exceed $600M run-rate materially
- Baseten aggregates GPU capacity across 18 cloud providers, creating structural dependency on third-party compute economics
Evidence Graph5 of 5 claims linked · 4 preserved sources
This graph uses only evidence frozen into the thesis at publication on 10/04/2026. “Retrieved source passage” is the preserved grounding excerpt the engine saw; “published note passage” is thesis context, not a source quote. Missing edges and dates remain visible.
Baseten's annualized revenue run-rate reached ~$600M by March 2026, up from ~$200M ARR in December 2025
QAI Finance reports ~$600M annualized run-rate March 2026 and ~$200M ARR Dec 2025
Retrieved source passage
Baseten Conviction Context B2B usage-based cloud platform. Sells 'deployments' rather than pure tokens: dedicated per-minute autoscaling GPU deployments, plus token-priced Model APIs and training/embeddings. Revenue scales with customer inference volume; COGS is dominated by rented/committed GPU capacity. Revenue ~$600M annualized run-rate March 2026); ~$200M ARR (Dec 2025 Rev growth ~20x YoY reported ~1,900%); annualized run-rate tripled Dec 2025 -> Mar 2026 Gross margin not disclosed inference peers such as Fireworks run ~50% per Sacra, well below the 70%+ SaaS norm because rented GPU capacity sits in COGS; treat as a peer-inference estimate, not a Baseten figure Op margin not disclosed
Baseten's Series F valuation of $13B implies ~21.7x annualized revenue, a premium that requires sustained 70%+ gross margins to justify against SaaS comps
Confirms $13B Series F valuation
Retrieved source passage
Announcing our Series F Announcing our Series F Baseten raised a $1.5B Series F and achieved a $13B valuation Authors Tuhin Srivastava Amir Haghighat Phil Howes Pankaj Gupta Last updated June 22, 2026 Today, we are thrilled to announce Baseten’s $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital, co-led by Sands Capital and Wellington Management, with participation from Battery Ventures, Blackbird, D.E. Shaw Ventures, Durable Capital Partners, Greylock, IVP, Verified Capital, and 01A. This is our fourth fundraise in 18 months, and we are grateful for our investors’ conviction and support and, most importantly, for the trust and partnership of our custo
Provides revenue figure enabling multiple calculation; notes peer inference margins ~50% vs 70%+ SaaS norm
Retrieved source passage
Baseten Conviction Context B2B usage-based cloud platform. Sells 'deployments' rather than pure tokens: dedicated per-minute autoscaling GPU deployments, plus token-priced Model APIs and training/embeddings. Revenue scales with customer inference volume; COGS is dominated by rented/committed GPU capacity. Revenue ~$600M annualized run-rate March 2026); ~$200M ARR (Dec 2025 Rev growth ~20x YoY reported ~1,900%); annualized run-rate tripled Dec 2025 -> Mar 2026 Gross margin not disclosed inference peers such as Fireworks run ~50% per Sacra, well below the 70%+ SaaS norm because rented GPU capacity sits in COGS; treat as a peer-inference estimate, not a Baseten figure Op margin not disclosed
Frames the central debate on whether inference remains differentiated or commoditizes
Retrieved source passage
Baseten — AI Inference Infrastructure Altis Research Know what’s really happening at Baseten before you sign on. Baseten is a managed inference platform that converts GPU capacity across 18 cloud providers into production-grade model serving for custom and open-weight AI workloads. With $2.1B raised and customers including Cursor, Notion, and Abridge, Baseten sits at the center of a critical debate: whether dedicated inference remains a differentiated, high-margin layer or commoditizes as compute scarcity eases. How we do it. Most "AI market intel" is a summarized crawl of press releases. This Altis research is built from our knowledge graph of sector expert calls, synthesized with pu
Inference peer gross margins run ~50% per Sacra data on Fireworks, well below the 70%+ SaaS benchmark, because rented GPU capacity sits in COGS
Explicitly cites ~50% peer inference margins and below-SaaS-norm explanation
Retrieved source passage
Baseten Conviction Context B2B usage-based cloud platform. Sells 'deployments' rather than pure tokens: dedicated per-minute autoscaling GPU deployments, plus token-priced Model APIs and training/embeddings. Revenue scales with customer inference volume; COGS is dominated by rented/committed GPU capacity. Revenue ~$600M annualized run-rate March 2026); ~$200M ARR (Dec 2025 Rev growth ~20x YoY reported ~1,900%); annualized run-rate tripled Dec 2025 -> Mar 2026 Gross margin not disclosed inference peers such as Fireworks run ~50% per Sacra, well below the 70%+ SaaS norm because rented GPU capacity sits in COGS; treat as a peer-inference estimate, not a Baseten figure Op margin not disclosed
Baseten surpassed its full-year CY26 revenue forecast by end of Q1 2026, implying forward revenue may exceed $600M run-rate materially
Apoorv Agrawal states Baseten surpassed full-year CY26 forecast by end of Q1 2026
Retrieved source passage
Why we are doubling down on Baseten - by Apoorv Agrawal Why we are doubling down on Baseten Apoorv Agrawal Jun 22, 2026 We backed Baseten in Q4 2025, and I wrote up the thesis then. Six months on, it has only gotten more obvious to us, and faster. By the end of Q1, Baseten had already surpassed the full-year CY26 forecast we had underwritten. The updated thesis is: - We believe inference will be one of the largest market in AI - Post-trained open source models deliver the best combination of capability, cost and control - Baseten provides the entire model supply chain to harness open source. Train, deploy, and serve models all on the same platform - Baseten = index on AI economy Inference
Baseten aggregates GPU capacity across 18 cloud providers, creating structural dependency on third-party compute economics
Altis Research describes 18 cloud provider integrations
Retrieved source passage
Baseten — AI Inference Infrastructure Altis Research Know what’s really happening at Baseten before you sign on. Baseten is a managed inference platform that converts GPU capacity across 18 cloud providers into production-grade model serving for custom and open-weight AI workloads. With $2.1B raised and customers including Cursor, Notion, and Abridge, Baseten sits at the center of a critical debate: whether dedicated inference remains a differentiated, high-margin layer or commoditizes as compute scarcity eases. How we do it. Most "AI market intel" is a summarized crawl of press releases. This Altis research is built from our knowledge graph of sector expert calls, synthesized with pu
What changed
Complete, timestamped thesis history.
-
Thesis publishedWATCH · LOW conviction
Setup
Baseten is a private, usage-based AI inference infrastructure platform that converts rented GPU capacity across 18 cloud providers into production-grade model serving for open-weight and custom AI workloads (G3). The company raised a $1.5B Series F at a $13B valuation in June 2026 — its fourth fundraise in 18 months — led by Altimeter, Conviction, and Spark, with co-leads Sands Capital and Wellington (G1). Total capital raised stands at $2.1B (G3). Customers include Cursor, Notion, and Abridge. The investment question is narrow but binary: does dedicated inference remain a differentiated, high-margin layer, or does it commoditize as GPU scarcity eases?
Evidence & Data
| Metric | Value | Source |
|---|---|---|
| Series F valuation | $13B | G1 |
| Series F raise | $1.5B | G1 |
| Total raised | $2.1B | G3 |
| Annualized run-rate (Mar 2026) | ~$600M | G4 |
| ARR (Dec 2025) | ~$200M | G4 |
| YoY growth | ~1,900% (~20x) | G4 |
| Gross margin | Not disclosed; peer ~50% | G4 |
| Revenue multiple at F | ~21.7x run-rate | Calculated: $13B / $600M |
The growth trajectory is extraordinary. Annualized run-rate tripled from ~$200M (Dec 2025) to ~$600M (Mar 2026), and an investor note claims Baseten had already surpassed its full-year CY26 forecast by end of Q1 (G2, G4). If the CY26 plan was even $400M ARR, surpassing it by Q1 implies the run-rate could be materially above $600M by now. But this is secondhand from a bull investor — treat with skepticism.
The central problem is margin structure. Baseten's gross margin is not disclosed. Peer inference platform Fireworks runs at ~50% gross margin per Sacra, well below the 70%+ SaaS norm, because rented or committed GPU capacity sits in COGS (G4). At a 21.7x revenue multiple, the market is pricing Baseten closer to a high-margin SaaS franchise than a low-margin compute reseller. If Baseten's margins resemble Fireworks (~50%), the multiple on gross profit is effectively ~43x — aggressive even for 20x growth. If margins approach 70%, the picture improves materially but still demands flawless execution.
Consensus view and variant perception. The consensus, reflected in the $13B valuation and blue-chid cap table, is that Baseten is the index play on AI inference — the picks-and-shovels beneficiary of open-weight model adoption, growing faster than any private peer, with platform lock-in via its full-stack train-deploy-serve workflow (G2). The market is likely underpricing margin compression risk. Inference is structurally a pass-through business when your COGS is someone else's GPU rental rate. As H100/H200 supply normalizes through 2026-2027, GPU rental rates should fall, and the spread Baseten captures between wholesale compute and retail serving will compress unless it builds genuine software differentiation (autoscaling, routing, post-training tooling) that commands a premium. The Altis research note frames this as the critical debate without resolving it (G3). No available data resolves it either.
Technical, sentiment, and macro stack. There is no public price chart, no technical indicator stack, and no analyst price-target consensus for a private company — I flag this explicitly rather than fabricate. News sentiment is strongly positive on growth metrics but silent on unit economics (G1, G2). Macro conditions matter: the AI infrastructure cycle is still in a capacity-constrained phase where GPU scarcity inflates inference pricing; a cyclical easing would be the primary macro risk vector. Rate environment is supportive of long-duration growth assets in 2026 but any tightening cycle would compress private valuation multiples directly.
Scenario Analysis
| Scenario | Probability | Price path | Thesis impact |
|---|---|---|---|
| Sustained scarcity + margin expansion to 60%+ | 25% | Next round at $18-22B; 30-40%+ upside from F | Bull case confirmed; inference is a moat |
| Status quo: 50% margins, growth slows to 5-8x | 40% | Flat to +15% at next round; multiple compresses as growth normalizes | WATCH thesis correct; no alpha |
| GPU glut: rental rates fall, margins compress to 35-40% | 25% | Down-round risk at $8-10B; 25-35% downside from F | Bear case; commoditization thesis |
| Platform disruption: hyperscaler native inference or open-source routing layer disintermediates | 10% | Severe down-round or distress; 50%+ downside | Tail risk; structural disintermediation |
EV = 0.25×(+32%) + 0.40×(+5%) + 0.25×(−30%) + 0.10×(−50%) = +8.0% − 7.5% − 5.0% = −4.5% expected return vs. Series F valuation, ±15% confidence band given private-market illiquidity and margin opacity.
The negative expected value is driven by the asymmetry: the bull case requires margin expansion that no data supports, while the bear cases require only the normalization of GPU supply that the hardware cycle practically guarantees over 12-24 months.
Catalysts & Risks
Catalysts (upside):
- Gross margin disclosure showing >55% — would re-rate the multiple instantly.
- Revenue run-rate exceeding $800M by Q3 2026 — confirms CY26 beat is accelerating, not a one-quarter artifact.
- Hyperscaler partnership or acquisition interest — Baseten's 18-provider aggregation layer is strategically valuable to a cloud platform lacking native inference tooling (G3).
Risks (downside):
- GPU supply normalization through 2026-2027 compresses rental spreads — the core thesis risk.
- Hyperscalers build competitive native inference services (AWS SageMaker, GCP Vertex) at lower margins, undercutting dedicated platforms.
- Open-source routing layers (vLLM, TGI improvements) reduce the value of managed inference orchestration.
- Customer concentration: Cursor and Notion are marquee but their inference spend is not disclosed; losing either would be material at this scale.
- $2.1B raised against ~$600M run-rate implies a burn rate that could accelerate if growth stalls (G3, G4).
What Changes Our Mind
This is a WATCH because the name is private (no T+30 cohort), the margin profile is undisclosed, and the expected value arithmetic is negative at the current valuation without a margin expansion catalyst. The falsifiable triggers that would move this to a directional call:
| Trigger | Threshold | Direction |
|---|---|---|
| Baseten discloses gross margin | >55% → BULLISH; <45% → BEARISH | Either |
| Next round valuation | >$18B confirms scarcity pricing power → BULLISH | Up |
| GPU rental rate index (H100 hourly) | Falls >30% from 2026 peak → BEARISH on margin compression | Down |
| Revenue run-rate stalls below $700M by Q3 2026 | Growth deceleration from 20x to <4x → BEARISH | Down |
| Hyperscaler launches native managed inference at <40% gross margin | Structural price war → BEARISH | Down |
The single most informative data point would be Baseten's actual gross margin. Until that surfaces, the $13B valuation is an option premium on margin expansion that the evidence does not support. We are watching, not betting.
Educational market commentary. Not investment advice. Private company analysis based on retrieved web sources cited inline; figures are unverified and may be incomplete.
More from QL Research
- RILLET · WATCH — RILLET (Private) — WATCH: Unicorn optics vs. revenue reality at $1B
- VRTX · WATCH — VRTX: WATCH — Multi-Franchise Inflection Priced In, Bearish Technicals Unconfirmed by Fundamentals
- BMY · WATCH — BMY: WATCH — Engine Says Buy, Track Record Says Not So Fast
- WONDERFUL · WATCH — WONDERFUL — Series C Valuation vs. Revenue Reality: A WATCH on 71x ARR
- ARCHITECT-LABS · WATCH — ARCHITECT-LABS (Private): AI-to-Silicon Thesis Premature — WATCH Pending A0 Validation
- AVGO · WATCH — AVGO: The $77 Gap Nobody Can Close — WATCH Pending AI-Capex Reacceleration Proof