← QL Research Library
WATCHBASETENLOW conviction · open✦ Commissioned

Baseten — Private AI Inference at $13B: Growth Priced, Margin Unknown

Published · entry price $ · machine-generated by the QuantLogix Thesis Engine and graded publicly at T+7/30/90 days · BASETEN charts & signals →

Baseten's $600M annualized run-rate (Mar 2026) and ~20x YoY growth make it the fastest-scaling private inference platform, but the $13B Series F valuation embeds 21.7x revenue at a margin profile that is almost certainly sub-50% given rented GPU COGS. The thesis turns entirely on whether dedicated inference sustains pricing power as GPU scarcity eases — a question no available data resolves. WATCH pending gross margin disclosure or a credible margin expansion signal.

★ Follow this thesis Get notified as the position plays out — claim triggers, T+7/30/90 grades, and invalidation, by push and in-app. Pro members can also opt into email updates.

Thesis

Evidence Graph5 of 5 claims linked · 4 preserved sources
Claim coverage5 / 5claims with evidence
Directional edges66 support · 0 challenge
Freshness3 / 4Latest dated source 06/29/2026
ConvictionLOW0 recorded changes

This graph uses only evidence frozen into the thesis at publication on 10/04/2026. “Retrieved source passage” is the preserved grounding excerpt the engine saw; “published note passage” is thesis context, not a source quote. Missing edges and dates remain visible.

C1
1 linked source

Baseten's annualized revenue run-rate reached ~$600M by March 2026, up from ~$200M ARR in December 2025

G4SupportsExternal sourceOpen source
Baseten (BASETEN) - AI model inference & deployment (MLOps… - QAI Finance
Publication date not preserved · qai.io

QAI Finance reports ~$600M annualized run-rate March 2026 and ~$200M ARR Dec 2025

Retrieved source passage
Baseten Conviction Context B2B usage-based cloud platform. Sells 'deployments' rather than pure tokens: dedicated per-minute autoscaling GPU deployments, plus token-priced Model APIs and training/embeddings. Revenue scales with customer inference volume; COGS is dominated by rented/committed GPU capacity. Revenue ~$600M annualized run-rate March 2026); ~$200M ARR (Dec 2025 Rev growth ~20x YoY reported ~1,900%); annualized run-rate tripled Dec 2025 -> Mar 2026 Gross margin not disclosed inference peers such as Fireworks run ~50% per Sacra, well below the 70%+ SaaS norm because rented GPU capacity sits in COGS; treat as a peer-inference estimate, not a Baseten figure Op margin not disclosed
C2Invalidation rule
3 linked sources

Baseten's Series F valuation of $13B implies ~21.7x annualized revenue, a premium that requires sustained 70%+ gross margins to justify against SaaS comps

G1SupportsExternal sourceOpen source
Announcing our Series F
Published · baseten.co

Confirms $13B Series F valuation

Retrieved source passage
Announcing our Series F Announcing our Series F Baseten raised a $1.5B Series F and achieved a $13B valuation Authors Tuhin Srivastava Amir Haghighat Phil Howes Pankaj Gupta Last updated June 22, 2026 Today, we are thrilled to announce Baseten’s $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital, co-led by Sands Capital and Wellington Management, with participation from Battery Ventures, Blackbird, D.E. Shaw Ventures, Durable Capital Partners, Greylock, IVP, Verified Capital, and 01A. This is our fourth fundraise in 18 months, and we are grateful for our investors’ conviction and support and, most importantly, for the trust and partnership of our custo
G4SupportsExternal sourceOpen source
Baseten (BASETEN) - AI model inference & deployment (MLOps… - QAI Finance
Publication date not preserved · qai.io

Provides revenue figure enabling multiple calculation; notes peer inference margins ~50% vs 70%+ SaaS norm

Retrieved source passage
Baseten Conviction Context B2B usage-based cloud platform. Sells 'deployments' rather than pure tokens: dedicated per-minute autoscaling GPU deployments, plus token-priced Model APIs and training/embeddings. Revenue scales with customer inference volume; COGS is dominated by rented/committed GPU capacity. Revenue ~$600M annualized run-rate March 2026); ~$200M ARR (Dec 2025 Rev growth ~20x YoY reported ~1,900%); annualized run-rate tripled Dec 2025 -> Mar 2026 Gross margin not disclosed inference peers such as Fireworks run ~50% per Sacra, well below the 70%+ SaaS norm because rented GPU capacity sits in COGS; treat as a peer-inference estimate, not a Baseten figure Op margin not disclosed
G3Context matchExternal sourceOpen source
Baseten — AI Inference Infrastructure Altis Research
Published · altis.vc

Frames the central debate on whether inference remains differentiated or commoditizes

Retrieved source passage
Baseten — AI Inference Infrastructure Altis Research Know what’s really happening at Baseten before you sign on. Baseten is a managed inference platform that converts GPU capacity across 18 cloud providers into production-grade model serving for custom and open-weight AI workloads. With $2.1B raised and customers including Cursor, Notion, and Abridge, Baseten sits at the center of a critical debate: whether dedicated inference remains a differentiated, high-margin layer or commoditizes as compute scarcity eases. How we do it. Most "AI market intel" is a summarized crawl of press releases. This Altis research is built from our knowledge graph of sector expert calls, synthesized with pu
C3Invalidation rule
1 linked source

Inference peer gross margins run ~50% per Sacra data on Fireworks, well below the 70%+ SaaS benchmark, because rented GPU capacity sits in COGS

G4SupportsExternal sourceOpen source
Baseten (BASETEN) - AI model inference & deployment (MLOps… - QAI Finance
Publication date not preserved · qai.io

Explicitly cites ~50% peer inference margins and below-SaaS-norm explanation

Retrieved source passage
Baseten Conviction Context B2B usage-based cloud platform. Sells 'deployments' rather than pure tokens: dedicated per-minute autoscaling GPU deployments, plus token-priced Model APIs and training/embeddings. Revenue scales with customer inference volume; COGS is dominated by rented/committed GPU capacity. Revenue ~$600M annualized run-rate March 2026); ~$200M ARR (Dec 2025 Rev growth ~20x YoY reported ~1,900%); annualized run-rate tripled Dec 2025 -> Mar 2026 Gross margin not disclosed inference peers such as Fireworks run ~50% per Sacra, well below the 70%+ SaaS norm because rented GPU capacity sits in COGS; treat as a peer-inference estimate, not a Baseten figure Op margin not disclosed
C4
1 linked source

Baseten surpassed its full-year CY26 revenue forecast by end of Q1 2026, implying forward revenue may exceed $600M run-rate materially

G2SupportsExternal sourceOpen source
Why we are doubling down on Baseten - by Apoorv Agrawal
Published · apoorv03.com

Apoorv Agrawal states Baseten surpassed full-year CY26 forecast by end of Q1 2026

Retrieved source passage
Why we are doubling down on Baseten - by Apoorv Agrawal Why we are doubling down on Baseten Apoorv Agrawal Jun 22, 2026 We backed Baseten in Q4 2025, and I wrote up the thesis then. Six months on, it has only gotten more obvious to us, and faster. By the end of Q1, Baseten had already surpassed the full-year CY26 forecast we had underwritten. The updated thesis is: - We believe inference will be one of the largest market in AI - Post-trained open source models deliver the best combination of capability, cost and control - Baseten provides the entire model supply chain to harness open source. Train, deploy, and serve models all on the same platform - Baseten = index on AI economy Inference
C5
1 linked source

Baseten aggregates GPU capacity across 18 cloud providers, creating structural dependency on third-party compute economics

G3SupportsExternal sourceOpen source
Baseten — AI Inference Infrastructure Altis Research
Published · altis.vc

Altis Research describes 18 cloud provider integrations

Retrieved source passage
Baseten — AI Inference Infrastructure Altis Research Know what’s really happening at Baseten before you sign on. Baseten is a managed inference platform that converts GPU capacity across 18 cloud providers into production-grade model serving for custom and open-weight AI workloads. With $2.1B raised and customers including Cursor, Notion, and Abridge, Baseten sits at the center of a critical debate: whether dedicated inference remains a differentiated, high-margin layer or commoditizes as compute scarcity eases. How we do it. Most "AI market intel" is a summarized crawl of press releases. This Altis research is built from our knowledge graph of sector expert calls, synthesized with pu

What changed

Complete, timestamped thesis history.

  1. Thesis published
    WATCH · LOW conviction

Setup

Baseten is a private, usage-based AI inference infrastructure platform that converts rented GPU capacity across 18 cloud providers into production-grade model serving for open-weight and custom AI workloads (G3). The company raised a $1.5B Series F at a $13B valuation in June 2026 — its fourth fundraise in 18 months — led by Altimeter, Conviction, and Spark, with co-leads Sands Capital and Wellington (G1). Total capital raised stands at $2.1B (G3). Customers include Cursor, Notion, and Abridge. The investment question is narrow but binary: does dedicated inference remain a differentiated, high-margin layer, or does it commoditize as GPU scarcity eases?

Evidence & Data

MetricValueSource
Series F valuation$13BG1
Series F raise$1.5BG1
Total raised$2.1BG3
Annualized run-rate (Mar 2026)~$600MG4
ARR (Dec 2025)~$200MG4
YoY growth~1,900% (~20x)G4
Gross marginNot disclosed; peer ~50%G4
Revenue multiple at F~21.7x run-rateCalculated: $13B / $600M

The growth trajectory is extraordinary. Annualized run-rate tripled from ~$200M (Dec 2025) to ~$600M (Mar 2026), and an investor note claims Baseten had already surpassed its full-year CY26 forecast by end of Q1 (G2, G4). If the CY26 plan was even $400M ARR, surpassing it by Q1 implies the run-rate could be materially above $600M by now. But this is secondhand from a bull investor — treat with skepticism.

The central problem is margin structure. Baseten's gross margin is not disclosed. Peer inference platform Fireworks runs at ~50% gross margin per Sacra, well below the 70%+ SaaS norm, because rented or committed GPU capacity sits in COGS (G4). At a 21.7x revenue multiple, the market is pricing Baseten closer to a high-margin SaaS franchise than a low-margin compute reseller. If Baseten's margins resemble Fireworks (~50%), the multiple on gross profit is effectively ~43x — aggressive even for 20x growth. If margins approach 70%, the picture improves materially but still demands flawless execution.

Consensus view and variant perception. The consensus, reflected in the $13B valuation and blue-chid cap table, is that Baseten is the index play on AI inference — the picks-and-shovels beneficiary of open-weight model adoption, growing faster than any private peer, with platform lock-in via its full-stack train-deploy-serve workflow (G2). The market is likely underpricing margin compression risk. Inference is structurally a pass-through business when your COGS is someone else's GPU rental rate. As H100/H200 supply normalizes through 2026-2027, GPU rental rates should fall, and the spread Baseten captures between wholesale compute and retail serving will compress unless it builds genuine software differentiation (autoscaling, routing, post-training tooling) that commands a premium. The Altis research note frames this as the critical debate without resolving it (G3). No available data resolves it either.

Technical, sentiment, and macro stack. There is no public price chart, no technical indicator stack, and no analyst price-target consensus for a private company — I flag this explicitly rather than fabricate. News sentiment is strongly positive on growth metrics but silent on unit economics (G1, G2). Macro conditions matter: the AI infrastructure cycle is still in a capacity-constrained phase where GPU scarcity inflates inference pricing; a cyclical easing would be the primary macro risk vector. Rate environment is supportive of long-duration growth assets in 2026 but any tightening cycle would compress private valuation multiples directly.

Scenario Analysis

ScenarioProbabilityPrice pathThesis impact
Sustained scarcity + margin expansion to 60%+25%Next round at $18-22B; 30-40%+ upside from FBull case confirmed; inference is a moat
Status quo: 50% margins, growth slows to 5-8x40%Flat to +15% at next round; multiple compresses as growth normalizesWATCH thesis correct; no alpha
GPU glut: rental rates fall, margins compress to 35-40%25%Down-round risk at $8-10B; 25-35% downside from FBear case; commoditization thesis
Platform disruption: hyperscaler native inference or open-source routing layer disintermediates10%Severe down-round or distress; 50%+ downsideTail risk; structural disintermediation
Scenario probabilities — engine-assigned odds, price paths on hover
Sustained scarcity + marg…25%Status quo: 50% margins, …40%GPU glut: rental rates fa…25%Platform disruption: hype…10%

EV = 0.25×(+32%) + 0.40×(+5%) + 0.25×(−30%) + 0.10×(−50%) = +8.0% − 7.5% − 5.0% = −4.5% expected return vs. Series F valuation, ±15% confidence band given private-market illiquidity and margin opacity.

The negative expected value is driven by the asymmetry: the bull case requires margin expansion that no data supports, while the bear cases require only the normalization of GPU supply that the hardware cycle practically guarantees over 12-24 months.

Catalysts & Risks

Catalysts (upside):

  • Gross margin disclosure showing >55% — would re-rate the multiple instantly.
  • Revenue run-rate exceeding $800M by Q3 2026 — confirms CY26 beat is accelerating, not a one-quarter artifact.
  • Hyperscaler partnership or acquisition interest — Baseten's 18-provider aggregation layer is strategically valuable to a cloud platform lacking native inference tooling (G3).

Risks (downside):

  • GPU supply normalization through 2026-2027 compresses rental spreads — the core thesis risk.
  • Hyperscalers build competitive native inference services (AWS SageMaker, GCP Vertex) at lower margins, undercutting dedicated platforms.
  • Open-source routing layers (vLLM, TGI improvements) reduce the value of managed inference orchestration.
  • Customer concentration: Cursor and Notion are marquee but their inference spend is not disclosed; losing either would be material at this scale.
  • $2.1B raised against ~$600M run-rate implies a burn rate that could accelerate if growth stalls (G3, G4).

What Changes Our Mind

This is a WATCH because the name is private (no T+30 cohort), the margin profile is undisclosed, and the expected value arithmetic is negative at the current valuation without a margin expansion catalyst. The falsifiable triggers that would move this to a directional call:

TriggerThresholdDirection
Baseten discloses gross margin>55% → BULLISH; <45% → BEARISHEither
Next round valuation>$18B confirms scarcity pricing power → BULLISHUp
GPU rental rate index (H100 hourly)Falls >30% from 2026 peak → BEARISH on margin compressionDown
Revenue run-rate stalls below $700M by Q3 2026Growth deceleration from 20x to <4x → BEARISHDown
Hyperscaler launches native managed inference at <40% gross marginStructural price war → BEARISHDown

The single most informative data point would be Baseten's actual gross margin. Until that surfaces, the $13B valuation is an option premium on margin expansion that the evidence does not support. We are watching, not betting.

Educational market commentary. Not investment advice. Private company analysis based on retrieved web sources cited inline; figures are unverified and may be incomplete.

More from QL Research

Share this research

Browse the full graded library →

QL Research is machine-generated educational market commentary, not investment advice. QuantLogix is not a registered investment adviser, broker-dealer, or financial planner. Theses, claims, verdicts, and grades are quantitative model outputs published for transparency and education; they are not recommendations to buy or sell any security. Markets involve substantial risk of loss. Past graded performance does not guarantee future results.