The AI Inference Stack

44 Companies 5 Categories AI Infrastructure

As AI moves from model development to widespread production usage, inference is becoming its own infrastructure stack. Performance increasingly depends not only on access to GPUs, but on specialized chips, compilers, runtimes, serving systems, and routing layers that optimize latency, throughput, energy consumption, hardware utilization, and cost per request.

Filter and ordering

Category

Inference chips & systems

8 companies

Companies designing specialized hardware and system architectures that make AI inference faster, more power-efficient, and more economical at scale.

d-Matrix

d-matrix.ai

d-Matrix builds AI inference compute hardware and software for data centers, including inference accelerators, networking, and an orchestration stack designed to make generative AI faster, more efficient, and cheaper to run at scale.

51-200 employees Employees Series C Stage 2019 Founded
Trajectory Maturity Level Proven

Confidence: High

Etched

etched.com

Etched designs frontier inference clusters—specialized AI chips, racks, software, and manufacturing methods intended to run frontier models with high throughput, low latency, and better power efficiency.

400+ employees Employees Series D+ Stage 2022 Founded
Trajectory Maturity Level Established

Confidence: High

Fractile

fractile.ai

Fractile designs AI inference chips, systems, and software intended to make frontier-model inference faster, more power-efficient, and lower-cost at scale.

90-100+ employees Employees Series B Stage 2022 Founded
Trajectory Maturity Level Emerging

Confidence: High

Groq

groq.com

Groq is a U.S.-based AI inference company that designs specialized LPU chips and builds GroqCloud and related systems to run AI models faster and more affordably at scale. ([groq.com](https://groq.com/newsroom/groq-raises-usd650m-to-scale-its-ai-inference-cloud-business?utm_source=openai))

51-200 employees Employees Series D+ Stage 2016 Founded
Trajectory Maturity Level Established

Confidence: High

OLIX

olix.com

OLIX is a photonic AI inference hardware company building the DX-1, a decode-focused accelerator and rack-scale system designed to improve inference throughput, latency, and energy efficiency.

about 130 employees Employees Series B Stage 2024 Founded
Trajectory Maturity Level Proven

Confidence: High

Positron

positron.ai

Positron builds purpose built hardware and systems for AI inference, including its Atlas appliance and upcoming Titan and Asimov silicon, with a focus on higher performance per dollar and lower power use. ([positron.ai](https://www.positron.ai/))

100+ Employees Series C Stage 2023 Founded
Trajectory Maturity Level Established

Confidence: High

Lamb Labs

lamb-labs.com

Lamb Labs builds model processing units and related software to make AI inference faster and more power efficient by hardcoding models into FPGA fabric and custom silicon.

2-10 employees Employees Pre Seed Stage 2026 Founded
Trajectory Maturity Level Early

Confidence: Medium

Volantis

volantissemi.ai

Volantis builds photonic AI inference hardware and system architecture meant to raise memory bandwidth and lower inference cost for very large models.

2-10 Employees Series A Stage 2022 Founded
Trajectory Maturity Level Early

Confidence: Medium

Category

Inference engines, compilers & runtimes

10 companies

Software that turns trained models into efficient executable workloads by adapting them to the underlying hardware and improving how inference is scheduled and run.

ZML

zml.ai

ZML builds a high-performance AI inference stack that helps run open-source models efficiently across different chips and production environments.

2-10 employees Employees Seed Stage 2023 Founded
Trajectory Maturity Level Early

Confidence: High

Modular

modular.com

Modular is an AI infrastructure company that builds compiler-driven software, including the MAX inference platform and Mojo language, to make model inference faster and more portable across heterogeneous hardware.

51-200 employees Employees Acquired Stage 2022 Founded
Acquired

Inferact

inferact.ai

Inferact is an AI infrastructure startup founded by the creators and core maintainers of vLLM, focused on making large-model inference cheaper, faster, and easier to serve at scale.

11-50 employees Employees Seed Stage 2025 Founded
Trajectory Maturity Level Emerging

Confidence: High

RadixArk

radixark.com

RadixArk is an infrastructure-first AI company that builds large-scale inference and training systems, including open-source serving and reinforcement-learning tooling, for developers, startups, enterprises, and research labs. ([radixark.com](https://www.radixark.com/))

11-50 employees Employees Seed Stage 2025 Founded
Trajectory Maturity Level Early

Confidence: High

Gimlet Labs

gimletlabs.ai

Gimlet Labs builds an inference cloud for agentic AI workloads, using heterogeneous hardware orchestration, a hardware-agnostic compiler, and kernel generation to make inference faster and more cost-efficient.

11-50 employees Employees Series A Stage 2025 Founded
Trajectory Maturity Level Emerging

Confidence: High

Infinity

infinity.inc

Infinity is an early-stage AI infrastructure company that uses AI to automatically generate, test, and optimize low-level inference code so non-NVIDIA chips can run models more efficiently.

2-10 employees Employees Seed Stage 2025 Founded
Trajectory Maturity Level Early

Confidence: High

RunAnywhere

runanywhere.ai

RunAnywhere is a research-first on-device AI inference platform that provides SDKs and runtimes to run multimodal models locally on iOS, Android, web, and edge devices with a control plane for deployment and policy management.

2-10 employees Employees Pre Seed Stage 2025 Founded
Trajectory Maturity Level Early

Confidence: Medium

Kog

kog.ai

Kog is a Paris-based AI infrastructure startup building a real-time inference engine for AI agents, with low-level GPU engineering and LLM architecture optimizations aimed at making token generation much faster on standard datacenter GPUs.

11 Employees Seed Stage 2023 Founded
Trajectory Maturity Level Early

Confidence: High

Callosum

callosum.com

Callosum builds software infrastructure for heterogeneous AI compute, with a focus on orchestrating inference across different chips and hardware stacks. ([callosum.com](https://www.callosum.com/?utm_source=openai))

20-30 Employees Seed Stage 2025 Founded
Trajectory Maturity Level Emerging

Confidence: High

Wafer

wafer.ai

Wafer builds AI agents and an inference platform that continuously optimize GPU serving stacks for faster, more efficient model inference.

2-10 employees Employees Series A Stage 2025 Founded
Trajectory Maturity Level Emerging

Confidence: High

Category

Model serving & deployment platforms

10 companies

Platforms that help teams deploy and operate models as reliable production services while handling scaling and infrastructure complexity.

Baseten

baseten.co

Baseten is an AI inference infrastructure platform that helps teams train, deploy, and serve models in production with autoscaling, observability, and optimized serving infrastructure.

201-500 employees Employees Series D+ Stage 2019 Founded
Trajectory Maturity Level Scaled

Confidence: High

Fireworks AI

fireworks.ai

Fireworks AI is an AI inference cloud and developer platform that helps teams deploy, serve, fine-tune, and optimize generative models in production with an emphasis on speed, cost efficiency, and control.

265 employees visible on LinkedIn Employees Series C Stage 2022 Founded
Trajectory Maturity Level Established

Confidence: High

BentoML

bentoml.com

BentoML is a unified inference platform and open-source framework for deploying, scaling, and operating AI models in production with infrastructure, autoscaling, observability, and cloud/on-prem deployment options. ([docs.bentoml.com](https://docs.bentoml.com/en/latest/?utm_source=openai))

51-200 employees Employees Acquired Stage 2019 Founded
Acquired

FriendliAI

friendli.ai

FriendliAI is a generative AI inference platform that helps teams deploy and operate large language and multimodal models in production with optimized speed, latency, and cost.

11-50 employees Employees Seed Stage 2021 Founded
Trajectory Maturity Level Emerging

Confidence: High

Modal

modal.com

Modal is a serverless cloud platform for AI/ML workloads that helps teams deploy and operate code, GPU jobs, inference, fine-tuning, and other production compute without managing infrastructure. ([modal.com](https://modal.com/company?utm_source=openai))

About 192 employees Employees Series C Stage 2021 Founded
Trajectory Maturity Level Established

Confidence: High

Replicate

replicate.com

A platform that lets developers run, fine-tune, and deploy machine-learning models through an API without managing underlying infrastructure.

26 employees Employees Acquired Stage 2019 Founded
Acquired

Together AI

together.ai

Together AI is a full-stack AI cloud platform that helps teams run, fine-tune, train, and deploy open-source models and AI agents with production-grade inference and infrastructure.

201-500 employees Employees Series C Stage 2022 Founded
Trajectory Maturity Level Established

Confidence: High

Runpod

runpod.io

Runpod is an AI developer cloud that lets teams build, train, fine-tune, deploy, and scale AI workloads on GPU infrastructure from one platform.

107 employees Employees Series A Stage 2022 Founded
Trajectory Maturity Level Proven

Confidence: High

OpenRelay

openrelay.inc

OpenRelay is a distributed, hardware-agnostic AI inference platform that connects consumer and datacenter GPUs into a fault-tolerant mesh for deploying production model workloads.

2 employees Employees Pre Seed Stage 2026 Founded
Trajectory Maturity Level Early

Confidence: High

Blackfuel

blackfuel.ai

Blackfuel is an AI inference company that provides dedicated infrastructure and an OpenAI compatible API for serving open weight models in production.

unknown Employees Unknown Stage 2026 Founded
Trajectory Maturity Level Proven

Confidence: High

Category

Inference gateways, routing & control

12 companies

Infrastructure that connects applications with model providers and provides a shared control layer for production inference.

OpenRouter

openrouter.ai

OpenRouter is an AI gateway and model marketplace that lets developers and enterprises route requests across hundreds of LLMs through a single API with built-in reliability, pricing, and fallback controls.

11-50 Employees Acquired Stage 2023 Founded
Acquired

Not Diamond

notdiamond.ai

Not Diamond builds an intelligent AI model router and prompt optimization platform that helps developers route requests across multiple models to improve quality while reducing inference cost and latency.

11-50 Employees Seed Stage 2023 Founded
Trajectory Maturity Level Emerging

Confidence: High

Portkey

portkey.ai

Portkey is an AI gateway and control-plane platform that gives developers and enterprises one unified layer for routing, observability, guardrails, and governance across multiple model providers.

11-50 employees Employees Acquired Stage 2023 Founded
Acquired

Helicone

helicone.ai

Helicone is an AI gateway and LLM observability platform for developers that provides a unified API plus logging, monitoring, routing, caching, and cost/performance analytics across model providers. ([ycombinator.com](https://www.ycombinator.com/companies/helicone?utm_source=openai))

5 Employees Acquired Stage 2023 Founded
Acquired

Requesty

requesty.ai

Requesty is an AI gateway and LLM router that gives teams a single OpenAI-compatible endpoint to access 600+ models with routing, fallbacks, caching, governance, observability, and spend controls. ([uk.linkedin.com](https://uk.linkedin.com/company/requesty))

11-50 Employees Seed Stage 2023 Founded
Trajectory Maturity Level Proven

Confidence: High

TrustedRouter

trustedrouter.com

TrustedRouter is an AI routing company that gives developers one OpenAI compatible API for models from multiple providers, with controls for price, reliability, region, and privacy posture.

unknown Employees Seed Stage 2026 Founded
Trajectory Maturity Level Early

Confidence: Medium

LiteLLM

litellm.ai

LiteLLM is an open source AI gateway and proxy that gives teams one interface to route, control, and monitor requests across many model providers.

10 Employees Seed Stage 2023 Founded
Trajectory Maturity Level Emerging

Confidence: High

TrueFoundry

truefoundry.com

TrueFoundry builds an enterprise AI gateway and deployment platform that gives teams a shared control layer for routing, governing, and observing model and agent traffic in production.

51-200 employees Employees Series A Stage 2021 Founded
Trajectory Maturity Level Proven

Confidence: High

Maxim AI

getmaxim.ai

Maxim AI builds an enterprise AI evaluation, observability, and gateway platform that helps teams route, govern, test, and monitor production AI applications.

11-50 Employees Seed Stage 2023 Founded
Trajectory Maturity Level Emerging

Confidence: High

Respan

respan.ai

Respan is an AI engineering platform for LLM and agent products that provides routing, observability, evaluations, and control for production inference. ([respan.ai](https://www.respan.ai/docs/documentation/overview?utm_source=openai))

20-25 employees Employees Seed Stage 2023 Founded
Trajectory Maturity Level Emerging

Confidence: High

Thesean

thesean.ai

Thesean is an AI inference research lab incubated by Martian. Its first product, Ship, is a best-execution endpoint that aims to reduce frontier-model inference cost while preserving model capability and behavior.

unknown Employees Unknown Stage 2026 Founded
Trajectory Maturity Level Early

Confidence: Medium

Concentrate AI

concentrate.ai

Concentrate AI builds an LLM gateway that gives teams one API for multiple model providers, with routing, failover, logging, spend controls, and access governance for production inference. ([concentrate.ai](https://concentrate.ai/?utm_source=openai))

2-10 employees Employees Pre Seed Stage 2025 Founded
Trajectory Maturity Level Early

Confidence: Medium

Category

On-device & local inference

4 companies

Infrastructure built to run AI models directly on user-owned or edge hardware rather than relying primarily on remote cloud inference. These platforms make local models easier to deploy and operate while accounting for the constraints of the underlying device.

Ollama

ollama.com

Ollama builds a local and cloud platform that makes it easy for developers and teams to run open AI models on their own hardware or nearby infrastructure.

14 Employees Series B Stage 2023 Founded
Trajectory Maturity Level Proven

Confidence: High

LM Studio

lmstudio.ai

LM Studio builds a desktop and developer platform for running open-source LLMs locally on user-owned hardware, with offline chat, model management, APIs, and tooling for app integration.

2-10 employees Employees Unknown Stage 2023 Founded
Trajectory Maturity Level Emerging

Confidence: High

Conifer

conifer.build

Conifer is an AI inference gateway that routes requests across local and cloud models, providing one interface, one account, and one bill for production inference.

4 Employees Pre Seed Stage 2025 Founded
Trajectory Maturity Level Nascent

Confidence: High

DiscreteStack

discretestack.com

DiscreteStack builds private AI infrastructure for enterprises so they can run open models on their own servers with flat rate pricing and full control over data and governance. It says the company is built in Europe and is designed for on premises deployment. ([discretestack.com](https://discretestack.com/))

2-10 Employees Seed Stage 2025 Founded
Trajectory Maturity Level Early

Confidence: High

Investors

Top-tier and active investors

Top-tier investors who invested

Did not invest

a16z crypto logo
Bessemer logo
Insight Partners logo
Balderton logo
Point Nine Capital logo

Other investors

Funding & Exits

Sep 2026

Volantis

Series A

$88M • Investors: Lachy Groom, Abstract Ventures, Sam Altman, Jeff Dean, Dylan Patel, John Doerr, VXI Capital, Triatomic, Susa Ventures, Dwarkesh Patel, Naveen Rao, Sholto Douglas

Volantis

Positron

Series C

$875M • Investors: NEA, Atreides Management, Valor Equity Partners, Andra Capital, Dylan Patel's SemiAnalysis Capital, Jim Clark

PR Newswire

Wafer

Series A

$40M • Investors: Marathon, Chemistry, Wing, AMD Ventures, Outset Capital, Fifty Years, Y Combinator, Jeff Dean, Guillermo Rauch, Andy Fang, Kyle Vogt, Akshay Kothari, Matthew Prince, Scott Stephenson

Wafer blog
Aug 2026

TrustedRouter

Seed

$1.2M • Investors: Sam Lessin / Slow Ventures, Bill Tai, Linda Avey, George Xing, Peter Livingston / Unpopular Ventures, Katelyn Donnelly / Avalanche VC, Gert Lanckriet, Holmes Wilson, Tory Reiss, Daniel Imberman, Michael Staton, Jason Fang, Capitoria Ventures, Alexey Komissarouk, Henri Roussez, others

TrustedRouter blog

Callosum

seed round

$100M • Investors: Atomico, Plural, DCVC, Sovereign AI, angel investors

tech.eu

OpenRouter

Acquisition

Acquirer: Stripe • Status: Announced

Stripe

Etched

Growth

$700M • Investors: Jane Street, Kleiner Perkins, Sequoia, Andreessen Horowitz, Peter Thiel, Tiger Global, Bain Capital Ventures, Neo, Stripes, Primary, Positive Sum, Diffusion, Argo, Blackstone

Etched

DiscreteStack

Seed

EUR 800K • Investors: CleverPine Ventures, Milen Manev, Stoil Vasilev, several smaller investors

DiscreteStack blog

OLIX

Series B

$312M • Investors: Hummingbird Ventures, Crane, Plural, Creandum, Phoenix Court, Transition, Fundomo, Arm, Hudson River Trading, Reed Hastings

www.finsmes.com
Jul 2026

Etched

Series C

$300M • Investors: Sequoia, Andreessen Horowitz, Jane Street, Diffusion, SK Hynix

Etched / GlobeNewswire

Infinity

Seed

$15M • Investors: Touring Capital, Principal Venture Partners

SiliconANGLE

ZML

Seed

$20M • Investors: 20VC, commit, AALVC, Drysdale Ventures, Kima Ventures, Kindred Capital VC, LocalGlobe, Puzzle Ventures

TechCrunch

Together AI

Series C

$800M • Investors: Aramco Ventures, NVIDIA, Vista Equity, General Catalyst, Emergence Capital, SE Ventures, Pegatron, Salesforce Ventures, March Capital, DTCP Growth, Lux Capital, Geodesic, PSP Partners

Together AI
Jun 2026

TrueFoundry

Acquisition

Acquirer: TrueFoundry • Status: Completed

press release

Runpod

Series A

$100M • Investors: Summit Partners

PR Newswire

Baseten

Growth

$1.5B • Investors: Altimeter Capital, Conviction Partners, Spark Capital, Sands Capital, Wellington Management, Battery Ventures, Blackbird, D.E. Shaw Ventures, Durable Capital Partners, Greylock, IVP, Verified Capital, 01A

Baseten

Groq

Growth capital

$650M • Investors: Disruptive, Infinitum

Groq
May 2026

Portkey

Acquisition

Deal: $140M • Acquirer: Palo Alto Networks • Status: Completed

Palo Alto Networks

OpenRouter

Series B

$113M • Investors: CapitalG, NVentures, ServiceNow Ventures, MongoDB Ventures, Snowflake Ventures, Databricks Ventures, AMP PBC, Pace Capital, Andreessen Horowitz, Menlo Ventures

Business Wire

Modal

Series C

$355M • Investors: General Catalyst, Redpoint Ventures, Menlo Ventures, Bain Capital Ventures, Accel, existing major investors

Modal

RadixArk

Seed

$100M • Investors: Accel, Spark Capital, NVentures, Salience Capital, A&E Investments, HOF Capital, Walden Catalyst Ventures, AMD, LDV Partners, WTT Investment, MediaTek

RadixArk blog

Fractile

Series B

$220M • Investors: Accel, Factorial Funds, Founders Fund, Conviction, Gigascale, 01A, Felicis, Buckley Ventures, 8VC

Fractile
Apr 2026

Portkey

Acquisition

Deal: $140M • Acquirer: Palo Alto Networks • Status: Announced

Palo Alto Networks

Wafer

Seed

$4M • Investors: Fifty Years, Liquid2, Y Combinator, Jeff Dean, Wojciech Zaremba, Arash Ferdowsi, Dan Fu

Official Wafer blog

LM Studio

Acquisition

Acquirer: LM Studio • Status: Announced

LM Studio blog

Callosum

government investment

Investors: UK Sovereign AI Fund

Callosum
Mar 2026

Respan

Seed

$5M • Investors: Gradient Ventures, Y Combinator, Hat-Trick Capital, XIAOXIAO FUND, Antigravity Capital, Alpen Capital

Respan official blog

Helicone

Acquisition

Acquirer: Mintlify • Status: Completed

Helicone Blog

Gimlet Labs

Series A

$80M • Investors: Menlo Ventures, Eclipse Ventures, Factory, Prosperity7, Triatomic

Gimlet Labs
Feb 2026

Portkey

Series A

$15M • Investors: Elevation Capital, Lightspeed

Portkey

OLIX

financing

$220M • Investors: Hummingbird Ventures, Plural, Vertex Ventures US, Entrepreneurs First, LocalGlobe

Cooley

BentoML

Acquisition

Acquirer: Modular • Status: Announced

BentoML blog
Jan 2026

Baseten

Growth

$300M • Investors: IVP, CapitalG, 01A, Altimeter, Battery Ventures, BOND, BoxGroup, Blackbird Ventures, Conviction, Greylock, NVIDIA

Baseten

Inferact

Seed

$150M • Investors: Andreessen Horowitz, Lightspeed Venture Partners, Sequoia Capital, Altimeter Capital, Redpoint Ventures, ZhenFund, The House Fund, Striker Venture Partners, Laude Ventures, Databricks Ventures, UC Berkeley Chancellor's Fund

Cooley

Etched

Series B

$500M • Investors: Stripes, Peter Thiel, Positive Sum, Ribbit Capital

Bloomberg
Nov 2025

d-Matrix

Series C

$275M • Investors: BullhoundCapital, Triatomic Capital, Temasek, Qatar Investment Authority, EDBI, M12, Nautilus Venture Partners, Industry Ventures, Mirae Asset

d-Matrix
Oct 2025

Fireworks AI

Series C

$250M • Investors: Lightspeed Venture Partners, Index Ventures, Evantic, Sequoia Capital

Fireworks AI Blog

Gimlet Labs

Seed

$12M • Investors: Factory, Lip-Bu Tan, Dylan Field, Rangarajan Raghuraman

Gimlet Labs
Sep 2025

Modal

Series B

$87M • Investors: Lux Capital, Amplify Partners, Redpoint Ventures

Modal

Baseten

Growth

$150M • Investors: BOND, Conviction, CapitalG, 01A, IVP, Spark, Greylock, Scribble Ventures, BoxGroup, Premji Invest

Baseten

Modular

Series C

$250M • Investors: US Innovative Technology Fund, DFJ Growth, GV, General Catalyst, Greylock Ventures

Modular

Groq

Growth

$750M • Investors: Disruptive, BlackRock, Neuberger Berman, DTCP, West Coast mutual fund manager, Samsung, Cisco, D1, Altimeter, 1789 Capital, Infinitum

Groq
Aug 2025

FriendliAI

Seed extension

$20M • Investors: Capstone Partners, Sierra Ventures, Alumni Ventures, KDB, KB Securities

FriendliAI
Jun 2025

Volantis

Seed

$9M • Investors: Alex Wang, Trevor Blackwell

Business Wire

Positron

Series A

$51.6M • Investors: Valor Equity Partners, Atreides Management, DFJ Growth, Flume Ventures, Resilience Reserve, 1517 Fund, Unless

LinkedIn
May 2025

Together AI

Acquisition

Acquirer: Together AI • Status: Announced

Together AI
Feb 2025

Positron

Seed

$23.5M • Investors: Flume Ventures, Valor Equity Partners, Atreides Management, Resilience Reserve

BusinessWire

TrueFoundry

Series A

$19M • Investors: Intel Capital, Peak XV Partners, Eniac Ventures, Jump Capital

TechCrunch

Baseten

Series C

$75M • Investors: IVP, Spark Capital, Greylock, Conviction, South Park Commons, Basecase, Lachy Groom, 01A

Baseten

Together AI

Series B

$305M • Investors: General Catalyst, Prosperity7, Salesforce Ventures, DAMAC Capital, NVIDIA, Kleiner Perkins, March Capital, Emergence Capital, Lux Capital, SE Ventures, Greycroft, Coatue, Definition, Cadenza Ventures, Long Journey Ventures, Brave Capital, Scott Banister, SK Telecom, John Chambers

Together AI
Dec 2024

OpenRouter

Series Seed

$10.8M

Forge

Together AI

Acquisition

Acquirer: Together AI • Status: Announced

Together AI
Aug 2024

Groq

Growth

$640M • Investors: BlackRock Private Equity Partners, Neuberger Berman, Type One Ventures, Cisco Investments, Global Brain's KDDI Open Innovation Fund III, Samsung Catalyst Fund

Groq
Jul 2024

Not Diamond

Seed

$2.3M • Investors: defy.vc, Jeff Dean, Julien Chaumond, Zack Kass, Ion Stoica, Tom Preston-Werner, Scott Belsky, Jeff Weiner

PR Newswire

Fireworks AI

Series B

$52M • Investors: Sequoia Capital, NVIDIA, AMD, MongoDB Ventures, Benchmark

Fireworks AI Blog

Fractile

Seed

$15M • Investors: Kindred Capital, NATO Innovation Fund, Oxford Science Enterprises

Data Center Dynamics
Jun 2024

Maxim AI

Seed

$3M • Investors: Elevation Capital, Undisclosed angel investors from Postman, Undisclosed angel investors from Chargebee, Undisclosed angel investors from Groww, Undisclosed angel investors from Razorpay, Undisclosed angel investors from Media.net

Maxim AI official blog

Etched

Series A

$120M • Investors: Primary Venture Partners, Positive Sum Ventures, Peter Thiel, Amjad Masad, Kyle Vogt

TechCrunch
May 2024

Runpod

Seed

$20M • Investors: Intel Capital, Dell Technologies Capital, Julien Chaummond, Nat Friedman, Adam Lewis

Business Wire
Mar 2024

Together AI

Series A

$106M • Investors: Salesforce Ventures, Coatue, Lux Capital, Kleiner Perkins, Emergence Capital, Prosperity7 Ventures, NEA, Greycroft, Definition Capital, Long Journey Ventures, Factory, Scott Banister, SV Angel, Clem Delangue, Soumith Chintala

Together AI

Fireworks AI

Series A

$25M • Investors: Benchmark, Sequoia Capital, Databricks Ventures, Frank Slootman, Alexandr Wang, Sheryl Sandberg, Howie Liu, Artisanal Ventures

FinSMEs

Baseten

Series B

$40M • Investors: IVP, Spark Capital, Greylock, South Park Commons, Lachy Groom, Base Case, Conviction

Baseten

Together AI

Series A extension

$106M • Investors: Salesforce Ventures, Coatue, Lux Capital, Kleiner Perkins, Emergence Capital, Prosperity7 Ventures, NEA, Greycroft, Definition Capital, Long Journey Ventures, Factory, SV Angel

Together AI blog
Dec 2023

Replicate

Series B

$40M • Investors: Andreessen Horowitz, NVentures, Heavybit, Sequoia Capital, Y Combinator

Replicate Blog
Nov 2023

Together AI

Series A

$102.5M • Investors: Kleiner Perkins, NVIDIA, Emergence Capital, NEA, Prosperity7 Ventures, Greycroft

TechCrunch
Oct 2023

Modal

Series A

$16M • Investors: Redpoint Ventures, Amplify Partners, Lux Capital, Definition Capital

Modal
Sep 2023

d-Matrix

Series B

$110M • Investors: Temasek, Playground Global, M12, SK Hynix, Nautilus Venture Partners, Entrada Ventures, Industry Ventures, Ericsson Ventures, Marlan Holding, Mirae Asset, Cortes Capital, Archerman Capital, TGC Square, Lam Capital, Samsung Ventures

d-Matrix
Aug 2023

Portkey

Seed

$3M • Investors: Lightspeed India, Dev Khare, Manjot Pahwa, Sanjeev Sisodiya, Adit Parekh, Ankit Gupta, Shyamal Hitesh Anadkat, Manish Jindal, Jake Seid, Oliver Jay, Aakrit Vaish, Pranay Gupta, Sandeep Krishnamurthy, Gaurav Mandlecha, Miten Sampat

Portkey

Modular

Series B

$100M • Investors: General Catalyst, GV, SV Angel, Greylock, Factory

Modular
Jun 2023

BentoML

Seed

$9M • Investors: DCM Ventures, Bow Capital, Firestreak Ventures

TechCrunch
May 2023

Together AI

Seed

$20M • Investors: Lux Capital, Factory, SV Angel, First Round Capital, Long Journey Ventures, A Capital, Robot Ventures, Common Metal, Definition Capital, Susa Ventures, Cadenza Ventures, SCB 10x, Scott Banister, Jeff Hammerbacher, Dawn Song, Alex Atallah, MC Lader, Lip-Bu Tan, Jakob Uszkoreit, Marc Bhargava, Jennifer Campbell, Chafic Kazoun, Sabrina Hahn, SongYee Yoon, Chase Lochmiller, Yi Sun, Dave Eisenberg, Panos Madamopoulos-Moraris, Zach Frankel

Together AI
Apr 2023

Positron

Seed

$6.5M • Investors: Jim Clark, strategic investors, additional investors

About Positron

Helicone

Pre-Seed

$1.5M • Investors: Coughdrop Capital, Y Combinator

CB Insights
Mar 2023

Etched

Seed

$5.4M • Investors: Primary Venture Partners, Positive Sum Ventures, Thomas Dohmke, Peter Thiel, Amjad Masad

Reuters / Yahoo Finance
Feb 2023

Replicate

Series A

$12.5M • Investors: Andreessen Horowitz, Y Combinator, Sequoia Capital, Dylan Field, Guillermo Rauch

Andreessen Horowitz
Jan 2023
Sep 2022

TrueFoundry

Seed

$2.3M • Investors: Sequoia India and Southeast Asia’s Surge, Eniac Ventures, Naval Ravikant, Dilip Khandelwal, Maneesh Sharma, Mike Boufford, Anthony Goldbloom

official blog
Jun 2022

Modular

Seed

$30M • Investors: GV, Greylock, The Factory, SV Angel

TechCrunch
Apr 2022

Baseten

Series A

$12M • Investors: Greylock, South Park Commons, Lachy Groom, Cristina Cordova, Dev Ittycheria, Jay Simons, Jean-Denis Greze

GlobeNewswire

Baseten

Seed

$8M • Investors: Greylock, South Park Commons Fund, AI Fund, Caffeinated Capital, Lachy Groom, Greg Brockman, Dylan Field, Mustafa Suleyman, DJ Patil

GlobeNewswire

d-Matrix

Series A

$44M • Investors: Playground Global, M12, SK Hynix, Nautilus Venture Partners, Marvell Technology, Entrada Ventures

Business Wire
Feb 2022
Apr 2021

Groq

Series C

$300M • Investors: Tiger Global Management, D1 Capital, The Spruce House Partnership, Addition, GCM Grosvenor, Xⁿ, Firebolt Ventures, General Global Capital, Tru Arrow Partners, TDK Ventures, XTX Ventures, Boardman Bay Capital Management, Infinitum Partners

Groq
May 2019
Sep 2018

Groq

Series A

$52.3M • Investors: Social Capital

TechCrunch
Jan 2017

On mobile, only the 10 most recent funding and exit groups are shown.

Activity

Recently added to this landscape

Last updated

Volantis builds photonic AI inference hardware and system architecture meant to raise memory bandwidth and lower inference cost for very large models.

Inference chips & systems

Blackfuel is an AI inference company that provides dedicated infrastructure and an OpenAI compatible API for serving open weight models in production.

Model serving & deployment platforms

Lamb Labs builds model processing units and related software to make AI inference faster and more power efficient by hardcoding models into FPGA fabric and custom silicon.

Inference chips & systems

Positron builds purpose built hardware and systems for AI inference, including its Atlas appliance and upcoming Titan and Asimov silicon, with a focus on higher performance per dollar and lower power use. ([positron.ai](https://www.positron.ai/))

Inference chips & systems

Concentrate AI builds an LLM gateway that gives teams one API for multiple model providers, with routing, failover, logging, spend controls, and access governance for production inference. ([concentrate.ai](https://concentrate.ai/?utm_source=openai))

Inference gateways, routing & control

Thesean is an AI inference research lab incubated by Martian. Its first product, Ship, is a best-execution endpoint that aims to reduce frontier-model inference cost while preserving model capability and behavior.

Inference gateways, routing & control

Respan is an AI engineering platform for LLM and agent products that provides routing, observability, evaluations, and control for production inference. ([respan.ai](https://www.respan.ai/docs/documentation/overview?utm_source=openai))

Inference gateways, routing & control

Maxim AI builds an enterprise AI evaluation, observability, and gateway platform that helps teams route, govern, test, and monitor production AI applications.

Inference gateways, routing & control

TrueFoundry builds an enterprise AI gateway and deployment platform that gives teams a shared control layer for routing, governing, and observing model and agent traffic in production.

Inference gateways, routing & control

LiteLLM is an open source AI gateway and proxy that gives teams one interface to route, control, and monitor requests across many model providers.

Inference gateways, routing & control

Wafer builds AI agents and an inference platform that continuously optimize GPU serving stacks for faster, more efficient model inference.

Inference engines, compilers & runtimes

TrustedRouter is an AI routing company that gives developers one OpenAI compatible API for models from multiple providers, with controls for price, reliability, region, and privacy posture.

Inference gateways, routing & control

DiscreteStack builds private AI infrastructure for enterprises so they can run open models on their own servers with flat rate pricing and full control over data and governance. It says the company is built in Europe and is designed for on premises deployment. ([discretestack.com](https://discretestack.com/))

On-device & local inference

Callosum builds software infrastructure for heterogeneous AI compute, with a focus on orchestrating inference across different chips and hardware stacks. ([callosum.com](https://www.callosum.com/?utm_source=openai))

Inference engines, compilers & runtimes

Requesty is an AI gateway and LLM router that gives teams a single OpenAI-compatible endpoint to access 600+ models with routing, fallbacks, caching, governance, observability, and spend controls. ([uk.linkedin.com](https://uk.linkedin.com/company/requesty))

Inference gateways, routing & control

Kog is a Paris-based AI infrastructure startup building a real-time inference engine for AI agents, with low-level GPU engineering and LLM architecture optimizations aimed at making token generation much faster on standard datacenter GPUs.

Inference engines, compilers & runtimes

RunAnywhere is a research-first on-device AI inference platform that provides SDKs and runtimes to run multimodal models locally on iOS, Android, web, and edge devices with a control plane for deployment and policy management.

Inference engines, compilers & runtimes

OLIX is a photonic AI inference hardware company building the DX-1, a decode-focused accelerator and rack-scale system designed to improve inference throughput, latency, and energy efficiency.

Inference chips & systems

Infinity is an early-stage AI infrastructure company that uses AI to automatically generate, test, and optimize low-level inference code so non-NVIDIA chips can run models more efficiently.

Inference engines, compilers & runtimes

OpenRelay is a distributed, hardware-agnostic AI inference platform that connects consumer and datacenter GPUs into a fault-tolerant mesh for deploying production model workloads.

Model serving & deployment platforms

Market analysis

The AI Inference Stack

A clear breakdown of the AI inference stack, from specialized chips and runtimes to model serving platforms and routing gateways, with key startup examples.

read more