The AI Inference Stack contains 44 companies across
5 categories.
As AI moves from model development to widespread production usage, inference is becoming its own infrastructure stack. Performance increasingly depends not only on access to GPUs, but on specialized chips, compilers, runtimes, serving systems, and routing layers that optimize latency, throughput, energy consumption, hardware utilization, and cost per request.
Inference chips & systems
Companies designing specialized hardware and system architectures that make AI inference faster, more power-efficient, and more economical at scale.
d-Matrix
Website: d-matrix.ai
d-Matrix builds AI inference compute hardware and software for data centers, including inference accelerators, networking, and an orchestration stack designed to make generative AI faster, more efficient, and cheaper to run at scale.
Category: Inference chips & systems.
Funding stage: Series C.
Founded: 2019.
Country: United States.
Employees: 51-200 employees.
Landscape fit: It fits the AI inference stack because its core products are purpose-built chips, rack-scale systems, and software for high-throughput, low-latency inference workloads rather than AI model training.
Funding rounds
-
Series C
announced 2025-11-01
: 275000000 USD
; investors: BullhoundCapital, Triatomic Capital, Temasek, Qatar Investment Authority, EDBI, M12, Nautilus Venture Partners, Industry Ventures, Mirae Asset
; source: d-Matrix
; https://www.d-matrix.ai/announcements/d-matrix-raises-275-million-to-power-the-age-of-ai-inference/
-
Series B
announced 2023-09-01
: 110000000 USD
; investors: Temasek, Playground Global, M12, SK Hynix, Nautilus Venture Partners, Entrada Ventures, Industry Ventures, Ericsson Ventures, Marlan Holding, Mirae Asset, Cortes Capital, Archerman Capital, TGC Square, Lam Capital, Samsung Ventures
; source: d-Matrix
; https://www.d-matrix.ai/announcements/d-matrix-announces-110-million-in-series-b-funding-to-make-generative-ai-commercially-viable-with-first-of-its-kind-inference-compute-platform/
-
Series A
announced 2022-04-01
: 44000000 USD
; investors: Playground Global, M12, SK Hynix, Nautilus Venture Partners, Marvell Technology, Entrada Ventures
; source: Business Wire
; https://www.businesswire.com/news/home/20220420005117/en/d-Matrix-Announces-%2444-Million-in-Funding-to-Build-a-One-of-a-Kind-Compute-Platform-Targeted-for-At-Scale-Transformer-AI-Datacenter-Inference
-
Series Seed
announced 2019-05-01
: 7260000 USD
; source: Forge Global
; https://forgeglobal.com/d-matrix_ipo/
Company timeline
-
2026-06-09:
Corsair entered full production
(product)
. d-Matrix announced Corsair was in full production and beginning volume shipments to priority hyperscalers, neoclouds, and frontier labs. ([d-matrix.ai](https://www.d-matrix.ai/announcements/d-matrix-corsair-ai-inference-platform-enters-full-production-to-meet-customer-demand/))
Source: d-Matrix.
https://www.d-matrix.ai/announcements/d-matrix-corsair-ai-inference-platform-enters-full-production-to-meet-customer-demand/
-
2026-06-09:
Moved from launch mode to deployment readiness after production demand surfaced
(gtm)
. d-Matrix announced that Corsair entered full production and would ship in volume to priority customers, emphasizing priority hyperscalers, neoclouds, and frontier labs.
Source: d-Matrix.
https://www.d-matrix.ai/announcements/d-matrix-corsair-ai-inference-platform-enters-full-production-to-meet-customer-demand/
-
2026-05-01:
Active hiring across ML, software, TPM, and hardware roles
(hr)
. d-Matrix had multiple live Ashby job postings for ML Research Intern, Principal LLM Inference Engineer, Staff Software Engineer for SIMD kernels, Principal AI/ML System Software Engineer, Principal Technical Program Manager, and Principal Systems Hardware Engineer.. Team size observed: 51-200 employees employees.
Source: Ashby job postings.
https://jobs.ashbyhq.com/d-Matrix
-
2026-04-02:
GigaIO data center business acquired to strengthen rack-scale deployment
(product)
. d-Matrix acquired GigaIO’s data center business and its core data-center technologies, including SuperNODE and the FabreX PCIe-based memory fabric, to support system-level deployments. ([d-matrix.ai](https://www.d-matrix.ai/announcements/acquisition-of-gigaio/))
Source: d-Matrix.
https://www.d-matrix.ai/announcements/acquisition-of-gigaio/
-
2026-03-12:
Gimlet partnership brought Corsair into heterogeneous inference clouds
(product)
. d-Matrix said Gimlet Labs would deploy Corsair accelerators alongside GPUs in Gimlet Cloud to improve latency and power efficiency for agentic workloads. ([d-matrix.ai](https://www.d-matrix.ai/announcements/gimlet/))
Source: d-Matrix.
https://www.d-matrix.ai/announcements/gimlet/
-
2025-11-18:
Deepened the inference stack with 3DIMC, Andes, and Alchip collaborations
(gtm)
. d-Matrix publicly tied its next-generation Raptor architecture to Andes CPU IP and Alchip 3D DRAM collaboration, reinforcing a memory-centric system architecture for future inference products.
Source: d-Matrix.
https://www.d-matrix.ai/announcements/d-matrix-and-alchip-announce-collaboration-on-worlds-first-3d-dram-solution-to-supercharge-ai-inference/
Etched
Website: etched.com
Etched designs frontier inference clusters—specialized AI chips, racks, software, and manufacturing methods intended to run frontier models with high throughput, low latency, and better power efficiency.
Category: Inference chips & systems.
Funding stage: Series D+.
Founded: 2022.
Country: United States.
Employees: 400+ employees.
Landscape fit: It fits the inference chips & systems category because its core product is specialized hardware and rack-scale infrastructure purpose-built to accelerate AI inference rather than general-purpose training.
Funding rounds
-
Series D
announced 2026-08-18
: 700000000 USD
; investors: Jane Street, Kleiner Perkins, Sequoia, Andreessen Horowitz, Peter Thiel, Tiger Global, Bain Capital Ventures, Neo, Stripes, Primary, Positive Sum, Diffusion, Argo, Blackstone
; source: Etched
; https://www.etched.com/progress/from-zero-to-one
-
Series C
announced 2026-07-23
: 300000000 USD
; investors: Sequoia, Andreessen Horowitz, Jane Street, Diffusion, SK Hynix
; source: Etched / GlobeNewswire
; https://www.globenewswire.com/news-release/2026/07/23/3332366/0/en/Etched-raises-300M-at-a-10-3B-Valuation-to-Scale-Production-of-Frontier-Scale-Inference-Hardware.html
-
Series B
announced 2026-01-01
: 500000000 USD
; investors: Stripes, Peter Thiel, Positive Sum, Ribbit Capital
; source: Bloomberg
; https://news.bloomberglaw.com/capital-markets/ai-chip-startup-etched-raises-500-million-to-take-on-nvidia
-
Series A
announced 2024-06-01
: 120000000 USD
; investors: Primary Venture Partners, Positive Sum Ventures, Peter Thiel, Amjad Masad, Kyle Vogt
; source: TechCrunch
; https://techcrunch.com/2024/06/25/etched-is-building-an-ai-chip-that-only-runs-transformer-models/
-
Seed
announced 2023-03-01
: 5400000 USD
; investors: Primary Venture Partners, Positive Sum Ventures, Thomas Dohmke, Peter Thiel, Amjad Masad
; source: Reuters / Yahoo Finance
; https://tech.yahoo.com/ai/articles/ai-startup-etched-raises-120-120324279.html
Company timeline
-
2026-08-18:
Etched announces first customer delivery to Jane Street
(milestone)
. Etched announced that it had shipped its first rack to Jane Street, its first customer, and that Jane Street was actively deploying the system into its workloads.
Source: Etched / GlobeNewswire.
https://www.globenewswire.com/news-release/2026/08/18/3347095/0/en/etched-raises-700m-at-a-21b-valuation-and-completes-first-customer-delivery-to-jane-street.html
-
2026-08-18:
Etched raises $700M at $21B valuation
(funding)
. Etched raised $700M in financing at a $21 billion valuation. The company said it will use the funds to scale production for customers.
Source: Etched.
https://www.etched.com/progress/from-zero-to-one
-
2026-07-23:
Etched raises $300M Series C at $10.3B valuation
(funding)
. Etched raised a $300M Series C led by Sequoia at a $10.3B valuation to scale production and customer deployments of its frontier inference systems.
Source: Etched / GlobeNewswire.
https://www.globenewswire.com/news-release/2026/07/23/3332366/0/en/Etched-raises-300M-at-a-10-3B-Valuation-to-Scale-Production-of-Frontier-Scale-Inference-Hardware.html
-
2026-07-10:
Careers page shows active multi-function hiring
(hr)
. Etched’s join page currently lists open roles across ASIC, platform, and software, with additional roles available behind a 'View More' control.. Team size observed: 400+ employees employees.
Source: Etched.
https://www.etched.com/join?trk=organization_guest_main-feed-card-text
-
2026-07-01:
Moved into production and manufacturing scale-up
(product)
. Etched said it had kicked off production for first racks, was working to fulfill over $1B in customer contracts, and had opened a Taiwan factory plus on-site data center, test house, and NPI lab. ([etched.com](https://www.etched.com/?utm_source=openai))
Source: Etched homepage.
https://www.etched.com/
-
2026-07-01:
Introduced Low Voltage Inference and Cluster Scale Memory
(product)
. Etched described two new system-level capabilities: Low Voltage Inference for higher sustained throughput and Cluster Scale Memory for lower-latency shared memory across chips. ([etched.com](https://www.etched.com/?utm_source=openai))
Source: Etched homepage.
https://www.etched.com/
Fractile
Website: fractile.ai
Fractile designs AI inference chips, systems, and software intended to make frontier-model inference faster, more power-efficient, and lower-cost at scale.
Category: Inference chips & systems.
Funding stage: Series B.
Founded: 2022.
Country: United Kingdom.
Employees: 90-100+ employees.
Landscape fit: It fits this category because it builds specialized inference hardware and system architecture aimed at reducing latency, power use, and cost for large-scale AI model serving.
Funding rounds
-
Series B
announced 2026-05-01
: 220000000 USD
; investors: Accel, Factorial Funds, Founders Fund, Conviction, Gigascale, 01A, Felicis, Buckley Ventures, 8VC
; source: Fractile
; https://www.fractile.ai/news/fractile-raises-220m-to-build-the-next-generation-of-inference-hardware
-
Seed
announced 2024-07-01
: 15000000 USD
; investors: Kindred Capital, NATO Innovation Fund, Oxford Science Enterprises
; source: Data Center Dynamics
; https://www.datacenterdynamics.com/en/news/uk-ai-chip-startup-fractile-emerges-from-stealth-with-15m-funding-round/
Company timeline
-
2026-07-01:
Employee base observed at roughly 90 to 100+ people
(hr)
. Two current job posts describe Fractile as a startup of approximately 90 people and separately as '100+ people' across London and Bristol.. Team size observed: 90-100+ employees employees.
Source: Fractile Greenhouse job posts.
https://job-boards.eu.greenhouse.io/fractile/jobs/4916990101
-
2026-07-01:
Broad hiring cluster appears across silicon, hardware, software, and operations
(hr)
. Fractile's public careers board shows 61-62 open roles spanning hardware test, hardware systems, manufacturing operations, finance & corporate operations, silicon, software, and programme management.. Team size observed: 90-100+ employees employees.
Source: Fractile Greenhouse careers page.
https://job-boards.greenhouse.io/Fractile
-
2026-05-01:
Refreshes website positioning around a new generation of processors
(product)
. Fractile’s website now describes the product as a new generation of processors with physically interleaved memory and compute, targeting thousands of tokens per second at scale.
Source: Fractile homepage.
https://www.fractile.ai/
-
2026-05-01:
Raises $220M to get first chips and systems to customers
(gtm)
. Fractile announced a $220M financing round and said the capital would accelerate getting its first chips and systems into customers’ hands, while hiring across the UK, US, and Taiwan. ([fractile.ai](https://www.fractile.ai/news/fractile-raises-220m-to-build-the-next-generation-of-inference-hardware?utm_source=openai))
Source: Fractile News / UK AI Hardware Plan.
https://www.fractile.ai/news/fractile-raises-220m-to-build-the-next-generation-of-inference-hardware
-
2026-05-01:
Expands product framing to first chips and systems for customers
(product)
. Fractile announced a $220M financing and said the capital would accelerate getting its first chips and systems into customers’ hands.
Source: Fractile news.
https://www.fractile.ai/news/fractile-raises-220m-to-build-the-next-generation-of-inference-hardware
-
2026-05-01:
Series B
(funding)
. Amount: 220000000 USD. Investors: Accel, Factorial Funds, Founders Fund, Conviction, Gigascale, 01A, Felicis, Buckley Ventures, 8VC
Source: Fractile.
https://www.fractile.ai/news/fractile-raises-220m-to-build-the-next-generation-of-inference-hardware
Groq
Website: groq.com
Groq is a U.S.-based AI inference company that designs specialized LPU chips and builds GroqCloud and related systems to run AI models faster and more affordably at scale. ([groq.com](https://groq.com/newsroom/groq-raises-usd650m-to-scale-its-ai-inference-cloud-business?utm_source=openai))
Category: Inference chips & systems.
Funding stage: Series D+.
Founded: 2016.
Country: United States.
Employees: 51-200 employees.
Landscape fit: It fits the inference chips & systems category because Groq builds purpose-built silicon and deployment infrastructure specifically for high-speed, energy-efficient AI inference rather than model training. ([groq.com](https://groq.com/newsroom/groq-raises-usd650m-to-scale-its-ai-inference-cloud-business?utm_source=openai))
Funding rounds
-
Growth capital
announced 2026-06-01
: 650000000 USD
; investors: Disruptive, Infinitum
; source: Groq
; https://groq.com/newsroom/groq-raises-usd650m-to-scale-its-ai-inference-cloud-business
-
Series E
announced 2025-09-01
: 750000000 USD
; investors: Disruptive, BlackRock, Neuberger Berman, DTCP, West Coast mutual fund manager, Samsung, Cisco, D1, Altimeter, 1789 Capital, Infinitum
; source: Groq
; https://groq.com/newsroom/groq-raises-750-million-as-inference-demand-surges
-
Series D
announced 2024-08-01
: 640000000 USD
; investors: BlackRock Private Equity Partners, Neuberger Berman, Type One Ventures, Cisco Investments, Global Brain's KDDI Open Innovation Fund III, Samsung Catalyst Fund
; source: Groq
; https://groq.com/newsroom/groq-raises-640m-to-meet-soaring-demand-for-fast-ai-inference
-
Series C
announced 2021-04-01
: 300000000 USD
; investors: Tiger Global Management, D1 Capital, The Spruce House Partnership, Addition, GCM Grosvenor, Xⁿ, Firebolt Ventures, General Global Capital, Tru Arrow Partners, TDK Ventures, XTX Ventures, Boardman Bay Capital Management, Infinitum Partners
; source: Groq
; https://groq.com/newsroom/groq-closes-300-million-fundraise
-
Series A
announced 2018-09-01
: 52300000 USD
; investors: Social Capital
; source: TechCrunch
; https://www.techcrunch.com/2018/09/05/secretive-semiconductor-startup-groq-raises-52m-from-social-capital/
Company timeline
-
2026-07-09:
Groq careers page shows open positions and internship recruiting
(hr)
. Groq’s careers page states that it has open positions, invites applicants to apply or share resumes, and says Winter 2026 internship applications will open in the fall.. Team size observed: 51-200 employees employees.
Source: Groq careers page.
https://groq.com/careers-at-groq?gh_jid=5950656003
-
2026-06-22:
Groq adds Alan Rice, Sinclair Schuller, and Rakesh Malhotra to leadership
(hr)
. Groq’s June 2026 funding announcement says Alan Rice is joining as COO and that Sinclair Schuller and Rakesh Malhotra are being appointed as CTO and CPO starting in July.. Team size observed: 51-200 employees employees.
Source: Groq newsroom.
https://groq.com/newsroom/groq-raises-usd650m-to-scale-its-ai-inference-cloud-business
-
2026-06-01:
Growth capital
(funding)
. Amount: 650000000 USD. Investors: Disruptive, Infinitum
Source: Groq.
https://groq.com/newsroom/groq-raises-usd650m-to-scale-its-ai-inference-cloud-business
-
2026-04-07:
Customer-story motion broadened to fintech and personal finance
(gtm)
. Groq published a Stash customer story showing Groq powering a real-time AI money coach for personal finance.
Source: Groq customer stories.
https://groq.com/customer-stories/how-stash-unlocked-the-future-of-personal-finance-with-groq
-
2025-11-25:
Groq launches MCP Connectors beta
(product)
. Groq introduced Groq-maintained MCP Connectors for Google Workspace services, reducing setup and infrastructure overhead for tool-using agents. ([groq.com](https://groq.com/blog/introducing-mcp-connectors-in-beta-on-groqcloud?utm_source=openai))
Source: Groq blog.
https://groq.com/blog/introducing-mcp-connectors-in-beta-on-groqcloud
-
2025-11-17:
Sydney APAC data center launched
(gtm)
. Groq announced its first Asia-Pacific infrastructure footprint in Sydney with Equinix to serve production inference workloads in the region.
Source: Groq newsroom.
https://groq.com/newsroom/groq-expands-to-asia-pacific-with-sydney-data-center-to-power-the-next-generation-of-ai-inference
OLIX
Website: olix.com
OLIX is a photonic AI inference hardware company building the DX-1, a decode-focused accelerator and rack-scale system designed to improve inference throughput, latency, and energy efficiency.
Category: Inference chips & systems.
Funding stage: Series B.
Founded: 2024.
Country: United Kingdom.
Employees: about 130 employees.
Landscape fit: It fits the inference chips & systems category because its core product is a specialized accelerator architecture for AI decode/inference, with explicit emphasis on co-design of logic, data movement, packaging, optics, and interconnect to lower cost and improve performance at scale.
Funding rounds
-
Series B
announced 2026-08-04
: 312000000 USD
; investors: Hummingbird Ventures, Crane, Plural, Creandum, Phoenix Court, Transition, Fundomo, Arm, Hudson River Trading, Reed Hastings
; source: www.finsmes.com
; https://www.finsmes.com/2026/08/olix-raises-312m-in-series-b-funding-at-a-3-3bn-valuation.html
-
financing
announced 2026-02-11
: 220000000 USD
; investors: Hummingbird Ventures, Plural, Vertex Ventures US, Entrepreneurs First, LocalGlobe
; source: Cooley
; https://www.cooley.com/news/coverage/2026/2026-02-11-olix-raises-%24220-million-in-financing
Company timeline
-
2026-08-05:
OLIX shows a current hiring cluster across multiple functions and geographies
(hr)
. The careers page and role pages indicate active hiring across operations, manufacturing, software, infrastructure, optics, ASIC, people, IT, mechanical, core tech, FPGA, finance, sourcing, and quality, with roles shown across London, Bristol, Austin, Toronto, San Francisco, the US, the UK, and Canada.
Source: Careers at OLIX.
https://olix.com/careers
-
2026-08-05:
Product snapshot
(product)
. OLIX is building the DX-1 decode accelerator and a rack-scale inference system using SRAM integrated with photonics and a co-designed stack spanning logic, data movement, packaging, optics, and interconnect.
Source: OLIX homepage.
https://olix.com
-
2026-08-05:
OLIX publicly details DX-1 decode accelerator and rack-scale production inference stack
(product)
. OLIX published its Compute Manifesto and multiple technical job posts describing DX-1 as the first accelerator architected specifically for decode, with rack-scale co-design across logic, data movement, packaging, optics, and interconnect. The software stack, hardware-in-the-loop testing, prototype-platform boards, mixed-signal verification, DFT, packaging, and laser/optics work show active productization and bring-up.
Source: OLIX official website.
https://olix.com/blog/compute-manifesto
-
2026-08-04:
Olix raises $312M Series B at $3.3B valuation
(funding)
. Olix raised $312 million in Series B financing at a $3.3 billion valuation. The round included existing investors Hummingbird Ventures, Crane, Plural, Creandum, Phoenix Court, and Transition, along with Fundomo, Arm, Hudson River Trading, and angel investor Reed Hastings.
Source: www.finsmes.com.
https://www.finsmes.com/2026/08/olix-raises-312m-in-series-b-funding-at-a-3-3bn-valuation.html
-
2026-08-01:
OLIX adds Nick McKeown board appointment and Matt Briers CFO hire
(hr)
. Alongside the Series B announcement, OLIX appointed Professor Nick McKeown to its board of directors and named Matt Briers as Chief Financial Officer.
Source: OLIX news.
https://olix.com/news/company-raises-series-b
-
2026-02-11:
Financing
(funding)
. Amount: USD 220M. Investors: Hummingbird Ventures, Plural, Vertex Ventures US, Entrepreneurs First, LocalGlobe
Source: Cooley.
https://www.cooley.com/news/coverage/2026/2026-02-11-olix-raises-%24220-million-in-financing
Positron
Website: positron.ai
Positron builds purpose built hardware and systems for AI inference, including its Atlas appliance and upcoming Titan and Asimov silicon, with a focus on higher performance per dollar and lower power use. ([positron.ai](https://www.positron.ai/))
Category: Inference chips & systems.
Funding stage: Series C.
Founded: 2023.
Country: United States.
Employees: 100+.
Landscape fit: It fits the inference chips and systems category because its core product is specialized hardware for running transformer models more efficiently than GPUs, and its roadmap centers on inference accelerator silicon and rack systems. ([positron.ai](https://www.positron.ai/))
Funding rounds
-
series_c
announced 2026-09-10
: 875000000 USD
; investors: NEA, Atreides Management, Valor Equity Partners, Andra Capital, Dylan Patel's SemiAnalysis Capital, Jim Clark
; source: PR Newswire
; https://www.prnewswire.com/news-releases/positron-ai-raises-875-million-at-a-5-billion-valuation-to-bring-its-next-generation-inference-silicon-to-market-302874601.html
-
series_b
announced 2026-02-01
: 230000000 USD
; source: LinkedIn
; https://www.linkedin.com/posts/positron-ai_were-excited-to-announce-our-230-million-activity-7425254724502949888-BFvF
-
series_a
announced 2025-06-01
: 51600000 USD
; investors: Valor Equity Partners, Atreides Management, DFJ Growth, Flume Ventures, Resilience Reserve, 1517 Fund, Unless
; source: LinkedIn
; https://www.linkedin.com/posts/mitesh7_i-am-very-excited-to-announce-that-positron-activity-7355640774505508864-xii6
-
seed
announced 2025-02-11
: 23500000 USD
; investors: Flume Ventures, Valor Equity Partners, Atreides Management, Resilience Reserve
; source: BusinessWire
; https://www.businesswire.com/news/home/20250211936199/en/Positron-Secures-%2423.5M-to-Design-And-Manufacture-Energy-Efficient-Made-In-America-AI-Chips
-
seed
announced 2023-04-01
: 6500000 USD
; investors: Jim Clark, strategic investors, additional investors
; source: About Positron
; https://www.positron.ai/about
Company timeline
-
2026-09-11:
HR snapshot
(hr)
. Positron says it has more than 100 employees across the United States, Canada, and Israel. The current hiring posture looks active, with many open roles in ASIC, silicon validation, power, software, and product.
Source: About Positron.
https://www.positron.ai/about
-
2026-09-11:
Product snapshot
(product)
. Positron's public product surface covers a production ready inference appliance, a 2027 rack scale inference system, custom accelerator silicon, and an OpenAI compatible inference API for deployment and testing.
Source: Positron home page.
https://www.positron.ai
-
2026-09-10:
Series C
(funding)
. Amount: USD 875M. Investors: NEA, Atreides Management, Valor Equity Partners, Andra Capital, Dylan Patel's SemiAnalysis Capital, Jim Clark
Source: PR Newswire.
https://www.prnewswire.com/news-releases/positron-ai-raises-875-million-at-a-5-billion-valuation-to-bring-its-next-generation-inference-silicon-to-market-302874601.html
-
2026-09-01:
Hiring cluster
(hr)
. Positron's careers page shows many active openings across ASIC, silicon, software, and product roles.
Source: Positron | Careers.
https://www.positron.ai/careers
-
2026-08-01:
Oracle deployment
(gtm)
. Positron said it had deployed more than 50 racks of Atlas at Oracle Cloud Infrastructure by August 2026.
Source: About Positron.
https://www.positron.ai/about
-
2026-02-01:
Series B
(funding)
. Amount: USD 230M
Source: LinkedIn.
https://www.linkedin.com/posts/positron-ai_were-excited-to-announce-our-230-million-activity-7425254724502949888-BFvF
Lamb Labs
Website: lamb-labs.com
Lamb Labs builds model processing units and related software to make AI inference faster and more power efficient by hardcoding models into FPGA fabric and custom silicon.
Category: Inference chips & systems.
Funding stage: Pre Seed.
Founded: 2026.
Country: United States.
Employees: 2-10 employees.
Landscape fit: It fits the inference chips and systems category because the company is explicitly building specialized chips, FPGA prototypes, and custom silicon to accelerate AI inference at lower power.
Company timeline
-
2026-09-15:
HR snapshot
(hr)
. Lamb Labs looks like a founder led startup with a small visible team and no clear public hiring signal.
Source: YC company page for Lamb Labs.
https://www.ycombinator.com/companies/lamb-labs/jobs
-
2026-09-15:
Product snapshot
(product)
. Lamb Labs builds MPUs, or model processing units, for faster and more power efficient AI inference. The company says it hardcodes the entire model into silicon and pairs the model with related software.
Source: Lamb Labs official website.
https://lamb-labs.com
-
2026-09-14:
YC names the founders
(hr)
. Y Combinator listed Niki Kotecha as CEO and Thomas Lanning as CTO for Lamb Labs.
Source: YC company page for Lamb Labs.
https://www.ycombinator.com/companies/lamb-labs/jobs
-
2026-08-01:
Launch coverage notes positioning
(gtm)
. Third party launch coverage described Lamb Labs as designing custom silicon for faster and more power efficient LLM inference.
Source: Imperium Magazine YC S26 launches post.
https://www.linkedin.com/posts/imperiummagazine_launches-of-the-week-slashy-built-by-harsha-activity-7490423692301340672-PU0p
-
2026-08-01:
FPGA prototype hits 6 W
(product)
. Thomas Lanning said the board was running a 7B parameter model at 6 W.
Source: Thomas Lanning FPGA performance post.
https://www.linkedin.com/posts/thomas-lanning_getting-the-board-warmed-up-only-6w-to-activity-7484439413893136385-Cdj7
-
2026-07-01:
Launch targets edge customers
(gtm)
. The launch post asked to speak with customers in robotics, humanoids, smartphones, wearables, smart devices, and industrial systems.
Source: Niki Kotecha launch post.
https://www.linkedin.com/posts/niki-kotecha_were-building-ultra-fast-ultra-low-power-activity-7475939701007290368-HmVc
Volantis
Website: volantissemi.ai
Volantis builds photonic AI inference hardware and system architecture meant to raise memory bandwidth and lower inference cost for very large models.
Category: Inference chips & systems.
Funding stage: Series A.
Founded: 2022.
Country: United States.
Employees: 2-10.
Landscape fit: It fits the inference chips and systems category because its product is a specialized hardware system for faster, more power efficient AI inference at scale.
Funding rounds
-
Series A
announced 2026-09-29
: 88000000 USD
; investors: Lachy Groom, Abstract Ventures, Sam Altman, Jeff Dean, Dylan Patel, John Doerr, VXI Capital, Triatomic, Susa Ventures, Dwarkesh Patel, Naveen Rao, Sholto Douglas
; source: Volantis
; https://volantissemi.ai/news-insights/our-88m-series-a-demolishing-the-memory-wall-with-photonics-post
-
Seed
announced 2025-06-12
: 9000000 USD
; investors: Alex Wang, Trevor Blackwell
; source: Business Wire
; https://finance.yahoo.com/news/volantis-unveils-photonic-compute-platform-134500548.html
Company timeline
-
2026-10-02:
HR snapshot
(hr)
. Volantis's Greenhouse board lists openings across advanced packaging, architecture, ASIC, photonics, PHY/SERDES, and systems. LinkedIn lists the company size as 2–10 employees.
Source: Volantis Greenhouse job board.
https://job-boards.greenhouse.io/volantissemiconductorinc
-
2026-10-02:
Product snapshot
(product)
. Volantis presents photonic interconnect and inference hardware intended to increase memory capacity and bandwidth for large-model datacenter inference. The site describes integrated micro-VCSELs and optical waveguides, and its product page presents A-1 as a rack-compatible system.
Source: Volantis home page.
https://volantissemi.ai
-
2026-10-01:
Volantis is hiring across core engineering
(hr)
. Volantis's job board shows an active hiring cluster across photonics, packaging, ASIC, architecture, and systems roles.
Source: Volantis Greenhouse job board.
https://job-boards.greenhouse.io/volantissemiconductorinc
-
2026-09-29:
Series A
(funding)
. Amount: USD 88M. Investors: Lachy Groom, Abstract Ventures, Sam Altman, Jeff Dean, Dylan Patel, John Doerr, VXI Capital, Triatomic, Susa Ventures, Dwarkesh Patel, Naveen Rao, Sholto Douglas
Source: Volantis.
https://volantissemi.ai/news-insights/our-88m-series-a-demolishing-the-memory-wall-with-photonics-post
-
2026-06-01:
Lewis Carpenter joins Volantis
(hr)
. Lewis Carpenter said in a LinkedIn post that he had joined Volantis Semi and would help bring the Photonic Motherboard to life.
Source: LinkedIn post by Lewis Carpenter.
https://www.linkedin.com/posts/lewis-carpenter-3271b777_excited-to-join-volantis-semi-and-help-bring-activity-7450973291021901824-1Bzs
-
2025-06-12:
Volantis emerges from stealth with photonic compute platform
(product)
. Volantis publicly unveiled its photonically integrated compute architecture for AI, built around direct laser modulation and wafer-scale integration to improve chip-to-chip communication efficiency.
Source: Business Wire.
https://finance.yahoo.com/news/volantis-unveils-photonic-compute-platform-134500548.html
Inference engines, compilers & runtimes
Software that turns trained models into efficient executable workloads by adapting them to the underlying hardware and improving how inference is scheduled and run.
ZML
Website: zml.ai
ZML builds a high-performance AI inference stack that helps run open-source models efficiently across different chips and production environments.
Category: Inference engines, compilers & runtimes.
Funding stage: Seed.
Founded: 2023.
Country: France.
Employees: 2-10 employees.
Landscape fit: It fits the AI inference stack category because its core product is inference-performance software and an LLM inference server designed to optimize how trained models execute on underlying hardware, including across multiple accelerator types.
Funding rounds
-
seed
announced 2026-07-01
: 20000000 USD
; investors: 20VC, commit, AALVC, Drysdale Ventures, Kima Ventures, Kindred Capital VC, LocalGlobe, Puzzle Ventures
; source: TechCrunch
; https://techcrunch.com/2026/07/08/hot-french-startup-zml-releases-free-product-to-speed-inference-across-lots-of-ai-chips/
Company timeline
-
2026-07-08:
LLMD alpha launch as a packaged serving product
(gtm)
. ZML launched LLMD, a self-contained LLM inference server with support for multiple open-source models, multi-accelerator execution, modern serving features, and OpenAI-compatible deployment.
Source: ZML Blog.
https://zml.ai/posts/llmd/
-
2026-07-08:
ZML/LLMD alpha launched as a universal LLM server
(product)
. ZML released LLMD alpha, a self-contained inference server supporting LLaMA, Gemma, Qwen, and Mistral across NVIDIA, AMD, TPU, Intel, and Apple Metal.
Source: ZML blog.
https://zml.ai/posts/llmd/
-
2026-07-01:
seed
(funding)
. Amount: 20000000 USD. Investors: 20VC, commit, AALVC, Drysdale Ventures, Kima Ventures, Kindred Capital VC, LocalGlobe, Puzzle Ventures
Source: TechCrunch.
https://techcrunch.com/2026/07/08/hot-french-startup-zml-releases-free-product-to-speed-inference-across-lots-of-ai-chips/
-
2026-04-07:
Tokenizer performance upgrade as a production-inference credibility signal
(gtm)
. ZML switched tokenizer handling to the IREE tokenizer and claimed up to roughly 10x tokenizer speedups in benchmarked cases.
Source: ZML Blog.
https://zml.ai/posts/iree-tokenizer/
-
2026-04-07:
Tokenization path switched for large latency gains
(product)
. ZML merged a tokenizer change from the Hugging Face Rust tokenizer crate to the IREE tokenizer for tokenizer.json models, claiming up to about 10x faster tokenization.
Source: ZML blog.
https://zml.ai/posts/iree-tokenizer/
-
2026-04-01:
Founder hiring mention on LinkedIn
(hr)
. Steeve Morin posted as founder of ZML and noted that the company was hiring.. Team size observed: 2-10 employees employees.
Source: Steeve Morin’s LinkedIn post.
https://www.linkedin.com/posts/steevemorin_introducing-zmlv2-zmlv2-is-a-complete-activity-7442262842008965120-hzMg
Modular
Website: modular.com
Modular is an AI infrastructure company that builds compiler-driven software, including the MAX inference platform and Mojo language, to make model inference faster and more portable across heterogeneous hardware.
Category: Inference engines, compilers & runtimes.
Funding stage: Acquired.
Founded: 2022.
Country: United States.
Employees: 51-200 employees.
Landscape fit: It fits the AI inference stack category because its core products optimize how trained models are compiled, scheduled, and run efficiently on underlying hardware, especially for inference workloads.
Funding rounds
-
Series C
announced 2025-09-01
: 250000000 USD
; investors: US Innovative Technology Fund, DFJ Growth, GV, General Catalyst, Greylock Ventures
; source: Modular
; https://www.modular.com/blog/modular-raises-250m-to-scale-ais-unified-compute-layer
-
Series B
announced 2023-08-01
: 100000000 USD
; investors: General Catalyst, GV, SV Angel, Greylock, Factory
; source: Modular
; https://www.modular.com/blog/weve-raised-100m-to-fix-ai-infrastructure-for-the-worlds-developers
-
Seed
announced 2022-06-01
: 30000000 USD
; investors: GV, Greylock, The Factory, SV Angel
; source: TechCrunch
; https://techcrunch.com/2022/06/30/modular-closes-30m-seed-round-to-simplify-the-process-of-developing-ai-systems/
Company timeline
-
2026-07-01:
Careers page shows 13 current openings
(hr)
. Modular’s careers page is actively recruiting, showing 13 openings and describing its interview flow, benefits, and onboarding process.. Team size observed: 51-200 employees employees.
Source: Modular Careers.
https://www.modular.com/company/careers
-
2026-07-01:
Active hiring cluster across engineering, product, and partnerships
(hr)
. Modular’s LinkedIn jobs page shows several openings posted within the past week, including Senior AI Kernel Engineer, Head of Hardware Partnerships, Senior AI Graph Compiler Engineer, Software Engineer, Hardware Enablement, and others.. Team size observed: 51-200 employees employees.
Source: LinkedIn.
https://www.linkedin.com/company/modular-ai/jobs
-
2026-06-24:
Qualcomm acquisition caps Modular's commercial evolution
(gtm)
. Qualcomm announced it would acquire Modular, citing Modular’s AI software platform as a foundation for generative and agentic AI across data center and edge environments. ([modular.com](https://www.modular.com/blog/qualcomm-to-acquire-modular?utm_source=openai))
Source: Modular.
https://www.modular.com/blog/qualcomm-to-acquire-modular
-
2026-06-18:
Modular 26.4 expands cloud motion around frontier models
(gtm)
. Modular 26.4 brought MoE serving to Modular Cloud, added support for new model architectures, and positioned Modular Cloud as a place to bring up frontier models quickly. ([modular.com](https://www.modular.com/blog/modular-26-4-sota-moe-serving-model-bringup-via-agent-skills-mojo-beta-2-and-more?utm_source=openai))
Source: Modular.
https://www.modular.com/blog/modular-26-4-sota-moe-serving-model-bringup-via-agent-skills-mojo-beta-2-and-more
-
2026-06-01:
Modular acquired by Qualcomm
(acquisition)
. Qualcomm announced it had reached an agreement to acquire Modular Inc., with the transaction expected to close in the second half of 2026 subject to customary closing conditions and regulatory approvals.. Acquirer: Qualcomm. Status: announced
Source: Qualcomm Investor Relations.
https://investor.qualcomm.com/news-events/press-releases/news-details/2026/Qualcomm-to-Acquire-Modular/default.aspx
-
2026-05-07:
MAX 26.3 added multi-GPU execution and video generation
(product)
. MAX 26.3 added video generation support for Wan 2.1/2.2, a distributed-aware Tensor API for multi-GPU model execution, and a new dedicated Mojo docs home at mojolang.org alongside Mojo 1.0 beta. ([docs.modular.com](https://docs.modular.com/max/changelog/))
Source: Modular changelog and forum announcement.
https://docs.modular.com/max/changelog/
Inferact
Website: inferact.ai
Inferact is an AI infrastructure startup founded by the creators and core maintainers of vLLM, focused on making large-model inference cheaper, faster, and easier to serve at scale.
Category: Inference engines, compilers & runtimes.
Funding stage: Seed.
Founded: 2025.
Country: United States.
Employees: 11-50 employees.
Landscape fit: It fits this category because it sits in the inference stack layer, optimizing how trained models run on underlying hardware and how inference is scheduled, served, and accelerated.
Funding rounds
-
seed
announced 2026-01-01
: 150000000 USD
; investors: Andreessen Horowitz, Lightspeed Venture Partners, Sequoia Capital, Altimeter Capital, Redpoint Ventures, ZhenFund, The House Fund, Striker Venture Partners, Laude Ventures, Databricks Ventures, UC Berkeley Chancellor's Fund
; source: Cooley
; https://www.cooley.com/news/coverage/2026/2026-01-22-inferact-announces-%24150-million-seed-financing
Company timeline
-
2026-07-02:
Partner ecosystem event with vLLM and Novita AI
(gtm)
. Inferact co-hosted a World’s Fair happy hour with vLLM and Novita AI, signaling ecosystem-based demand generation and relationship building around open models and inference infrastructure.
Source: Inferact LinkedIn company page.
https://www.linkedin.com/company/inferact
-
2026-07-01:
Current technical hiring push
(hr)
. Inferact is publicly recruiting engineers and researchers, with visible openings across inference, performance and scale, cloud orchestration, and exceptional generalist roles.
Source: Inferact / Ashby job postings.
https://jobs.ashbyhq.com/inferact/43c0ca54-fcf5-41fa-83a1-38800c75ccc0
-
2026-06-02:
Co-marketed enterprise inference optimization with DigitalOcean
(gtm)
. A June 2026 DigitalOcean blog post co-authored by Simon Mo described prefix-aware routing and prefix caching work built with large customers, positioning Inferact as a partner that helps ship optimizations from the engine layer into production offerings.
Source: DigitalOcean blog.
https://www.digitalocean.com/blog/reduce-llm-inference-costs-prefix-caching
-
2026-05-01:
Conference and community presence around open-source inference
(gtm)
. Inferact showed up at MLSys 2026 with a booth and sponsored lightning talk, using the event to reinforce its technical authority in open-source inference and to meet infrastructure practitioners directly.
Source: Inferact LinkedIn company page.
https://www.linkedin.com/company/inferact
-
2026-01-22:
Thought-leadership amplification through a16z podcast
(gtm)
. Inferact’s founders used an a16z podcast to explain the company’s thesis: inference is becoming one of the hardest problems in AI infrastructure and the company is building a universal open-source inference layer.
Source: a16z podcast.
https://a16z.com/podcast/inferact-building-the-infrastructure-that-runs-modern-ai/
-
2026-01-13:
Customer-performance proof through DigitalOcean and Character.ai
(gtm)
. DigitalOcean announced that its inference platform delivered 2x production inference throughput and 50% lower cost per token for Character.ai, citing optimization work involving vLLM and Inferact.
Source: DigitalOcean press release.
https://investors.digitalocean.com/news/news-details/2026/DigitalOceans-Inference-Cloud-Platform-Powered-by-AMD-Instinct-GPUs-Delivers-2X-Production-Inference-Performance-for-Character-ai/default.aspx
RadixArk
Website: radixark.com
RadixArk is an infrastructure-first AI company that builds large-scale inference and training systems, including open-source serving and reinforcement-learning tooling, for developers, startups, enterprises, and research labs. ([radixark.com](https://www.radixark.com/))
Category: Inference engines, compilers & runtimes.
Funding stage: Seed.
Founded: 2025.
Country: United States.
Employees: 11-50 employees.
Landscape fit: It fits the AI inference stack because RadixArk explicitly builds and commercializes inference engines, compilers/schedulers-adjacent infrastructure, and serving/training systems designed to make model execution faster, cheaper, and more reliable. ([radixark.com](https://www.radixark.com/))
Funding rounds
-
Seed
announced 2026-05-01
: 100000000 USD
; investors: Accel, Spark Capital, NVentures, Salience Capital, A&E Investments, HOF Capital, Walden Catalyst Ventures, AMD, LDV Partners, WTT Investment, MediaTek
; source: RadixArk blog
; https://www.radixark.com/blog/radixark-launches-100m-seed
Company timeline
-
2026-07-01:
Recent design hire visible on company LinkedIn
(hr)
. A LinkedIn update on the RadixArk company page says a designer joined the company.. Team size observed: 11-50 employees employees.
Source: RadixArk LinkedIn.
https://www.linkedin.com/company/radixark
-
2026-07-01:
Open hiring across multiple functions
(hr)
. RadixArk's careers page shows 23 open roles spanning engineering, legal, operations, and sales, including leadership and technical positions.. Team size observed: 11-50 employees employees.
Source: RadixArk Greenhouse.
https://job-boards.greenhouse.io/radixark
-
2026-05-05:
Official company launch with managed AI infrastructure platform
(product)
. RadixArk publicly launched with a $100M seed announcement and described its product as managed infrastructure and tooling built on SGLang for inference and Miles for reinforcement learning/post-training.
Source: RadixArk blog.
https://www.radixark.com/blog/radixark-launches-100m-seed
-
2026-05-05:
RadixArk launch included active hiring
(hr)
. RadixArk officially launched and said it was hiring for its core team while inviting candidates to send resumes to the company.. Team size observed: 11-50 employees employees.
Source: RadixArk Blog.
https://www.radixark.com/blog/radixark-launches-100m-seed
-
2026-05-05:
Public launch of RadixArk with seed-funded open-infrastructure positioning
(gtm)
. RadixArk officially launched with a public seed announcement and a company narrative centered on open, frontier AI infrastructure built on SGLang and Miles, plus managed infrastructure for teams building at scale.
Source: RadixArk Blog.
https://www.radixark.com/blog/radixark-launches-100m-seed
-
2026-05-01:
Commercial platform terms and free-tier packaging introduced
(gtm)
. Updated terms describe a fee-based platform with a free tier for non-commercial use, registration requirements, and paid access governed by a pricing page.
Source: RadixArk Terms of Service.
https://www.radixark.com/tos
Gimlet Labs
Website: gimletlabs.ai
Gimlet Labs builds an inference cloud for agentic AI workloads, using heterogeneous hardware orchestration, a hardware-agnostic compiler, and kernel generation to make inference faster and more cost-efficient.
Category: Inference engines, compilers & runtimes.
Funding stage: Series A.
Founded: 2025.
Country: United States.
Employees: 11-50 employees.
Landscape fit: It fits the AI inference stack category because its core product optimizes how trained models are compiled, scheduled, and executed across different accelerators to improve inference efficiency.
Funding rounds
-
series A
announced 2026-03-01
: 80000000 USD
; investors: Menlo Ventures, Eclipse Ventures, Factory, Prosperity7, Triatomic
; source: Gimlet Labs
; https://gimletlabs.ai/blog/announcing-series-a
-
seed
announced 2025-10-01
: 12000000 USD
; investors: Factory, Lip-Bu Tan, Dylan Field, Rangarajan Raghuraman
; source: Gimlet Labs
; https://gimletlabs.ai/blog/introducing-gimlet-labs
Company timeline
-
2026-06-29:
MLCommons membership announced for agentic inference benchmarks
(product)
. Gimlet said it joined MLCommons to help establish open benchmarks for agentic inference. ([gimletlabs.ai](https://gimletlabs.ai/blog/tags/gimlet-labs?utm_source=openai))
Source: Gimlet Labs Blog.
https://gimletlabs.ai/blog/building-new-benchmarks-for-the-agentic-era-with-mlcommons
-
2026-06-29:
Joined MLCommons to push agentic inference benchmarks
(gtm)
. Gimlet said it became a member of MLCommons and would help establish open benchmarks for agentic inference.
Source: Gimlet Labs blog.
https://gimletlabs.ai/blog
-
2026-06-01:
Hiring cluster expanded across core functions
(hr)
. Gimlet Labs publicly listed openings and recruitment needs across infrastructure, compilers, distributed systems, ML research, recruiting, legal, and marketing.. Team size observed: 11-50 employees employees.
Source: Ashby / LinkedIn / company site.
https://jobs.ashbyhq.com/gimlet
-
2026-05-01:
Head of Talent Acquisition joined
(hr)
. Gimlet Labs publicly welcomed Elliot as Head of Talent Acquisition.. Team size observed: 11-50 employees employees.
Source: LinkedIn.
https://www.linkedin.com/posts/gimletlabs_we-are-thrilled-to-welcome-elliot-as-our-activity-7457153336996167680-inDm
-
2026-03-23:
Series A announcement highlighted customer traction and enterprise credibility
(gtm)
. Gimlet announced its $80M Series A and said its customer base had tripled in five months, including a top frontier lab and a hyperscaler.
Source: Gimlet Labs blog.
https://gimletlabs.ai/blog/announcing-series-a
-
2026-03-11:
Performance proof through heterogeneous hardware and d-Matrix work
(gtm)
. Gimlet published a sequence of technical posts showing how it routes inference stages to the best-suited hardware, including a d-Matrix Corsair analysis with 2-10X latency improvement claims.
Source: Gimlet Labs blog.
https://gimletlabs.ai/blog/low-latency-spec-decode-corsair
Infinity
Website: infinity.inc
Infinity is an early-stage AI infrastructure company that uses AI to automatically generate, test, and optimize low-level inference code so non-NVIDIA chips can run models more efficiently.
Category: Inference engines, compilers & runtimes.
Funding stage: Seed.
Founded: 2025.
Country: United States.
Employees: 2-10 employees.
Landscape fit: It fits the AI inference stack because it builds the software layer that optimizes model execution on hardware, including kernel generation and inference runtime performance for AI chips.
Funding rounds
-
Seed
announced 2026-07-01
: 15000000 USD
; investors: Touring Capital, Principal Venture Partners
; source: SiliconANGLE
; https://siliconangle.com/2026/07/20/infinity-raises-15m-run-ai-inference-chipset/
Company timeline
-
2026-07-21:
Careers page shows active hiring across 7 open roles
(hr)
. Infinity’s careers page was updated to show “We’re Hiring” with 7 open roles, including 6 engineering roles, indicating an active hiring push at the early-stage company.
Source: Infinity careers page.
https://infinity.inc/careers
-
2026-07-01:
Dual-sided partner motion for hardware and inference providers
(gtm)
. Infinity’s partnership pages formalized two commercial paths: hardware enablement for chip vendors and optimized serving stacks for cloud and inference providers.
Source: Infinity partnerships pages.
https://infinity.inc/partnerships
-
2026-07-01:
Seed
(funding)
. Amount: 15000000 USD. Investors: Touring Capital, Principal Venture Partners
Source: SiliconANGLE.
https://siliconangle.com/2026/07/20/infinity-raises-15m-run-ai-inference-chipset/
-
2026-04-01:
d-Matrix deployment announcement
(product)
. Infinity announced it had brought Qwen3 up on new silicon in days, with matrix multiplications at 90%+ of theoretical peak within 10 hours and end-to-end model execution within 10 days.
Source: Infinity Research.
https://infinity.inc/research
-
2026-04-01:
Design-partner validation on d-Matrix silicon
(gtm)
. Infinity announced that it brought Qwen3 running end-to-end on d-Matrix silicon in about 10 days and reached over 90% of theoretical peak on matrix multiplication within 10 hours.
Source: Infinity research page.
https://infinity.inc/research
-
2026-03-01:
Qwen3 inference stack case study released
(product)
. Infinity released a case study showing its infy system wrote an inference engine from scratch and optimized Qwen3-8B to beat vLLM on throughput.
Source: Infinity Case Study.
https://infinity.inc/case-studies/qwen3-optimization
RunAnywhere
Website: runanywhere.ai
RunAnywhere is a research-first on-device AI inference platform that provides SDKs and runtimes to run multimodal models locally on iOS, Android, web, and edge devices with a control plane for deployment and policy management.
Category: Inference engines, compilers & runtimes.
Funding stage: Pre Seed.
Founded: 2025.
Country: United States.
Employees: 2-10 employees.
Landscape fit: It fits the AI inference stack because it builds the runtime/inference layer that executes trained models efficiently on consumer hardware, including custom kernels, SDKs, and scheduling/control-plane software for local inference.
Funding rounds
-
exempt_offering
announced 2025-10-16
; source: SEC Form D — RunAnywhere, Inc.
; https://www.sec.gov/Archives/edgar/data/2092245/000209224525000001/xslFormDX01/primary_doc.xml
Company timeline
-
2026-08-10:
Product snapshot
(product)
. Production-grade on-device AI platform with SDKs and runtimes for local multimodal inference on iOS, Android, web, macOS, and edge devices, plus a control plane for model deployment, routing, policy management, OTA updates, observability, and fleet operations.
Source: RunAnywhere Documentation.
https://docs.runanywhere.ai/index
-
2026-08-05:
RunAnywhere Android app distribution and usage milestone
(product)
. The Google Play listing shows the Android app is publicly available, had 1k+ downloads, and supports local chat, vision, document QA, voice, tools, and Qualcomm Hexagon NPU acceleration.
Source: Google Play listing.
https://play.google.com/store/apps/details?hl=en_GB&id=com.runanywhere.runanywhereai
-
2026-08-05:
RunAnywhere Android app distribution milestone
(gtm)
. The Google Play listing shows RunAnywhere’s Android app is publicly distributed, with 1k+ downloads and support for multimodal on-device AI features.
Source: Google Play listing.
https://play.google.com/store/apps/details?hl=en_GB&id=com.runanywhere.runanywhereai
-
2026-07-19:
PrismML Bonsai 1-bit models live in RunAnywhere apps
(product)
. RunAnywhere said PrismML’s 1-bit Bonsai models were live in its iOS, Android, and macOS apps, including a 27B 1-bit model on phone and first true 1-bit model on an NPU.
Source: RunAnywhere Official Blog.
https://www.runanywhere.ai/blog/bonsai-27b-1-bit-models-on-phone
-
2026-07-14:
LLM Hub iOS port using RunAnywhere SDK
(gtm)
. An indie developer said he ported the local-AI app LLM Hub to iOS in six weeks with RunAnywhere SDK, and the app was already doing about $2,000 MRR before the port.
Source: RunAnywhere Official Blog.
https://www.runanywhere.ai/blog/llm-hub-ios-case-study
-
2026-06-25:
QHexRT live for Qualcomm Hexagon NPU inference
(product)
. RunAnywhere’s research page states QHexRT launched on 2026-06-25 as full-stack NPU inference for Qualcomm Hexagon devices.
Source: RunAnywhere Research.
https://www.runanywhere.ai/research
Kog
Website: kog.ai
Kog is a Paris-based AI infrastructure startup building a real-time inference engine for AI agents, with low-level GPU engineering and LLM architecture optimizations aimed at making token generation much faster on standard datacenter GPUs.
Category: Inference engines, compilers & runtimes.
Funding stage: Seed.
Founded: 2023.
Country: France.
Employees: 11.
Landscape fit: It fits the inference engines, compilers & runtimes category because its core product is the Kog Inference Engine, which co-designs model architecture, runtime scheduling, and GPU kernels to accelerate inference and reduce decoding latency.
Funding rounds
-
seed
announced 2026-05-28
: 5000000 USD
; investors: Varsity VC, Bpifrance Deep Tech Program
; source: Official blog / company announcement
; https://blog.kog.ai/real-time-llm-inference-on-standard-gpus-3-000-tokens-s-per-request
Company timeline
-
2026-08-15:
Product snapshot
(product)
. Kog currently presents a real-time inference stack for AI agents, centered on a low-latency engine, GPU-kernel optimization, and a latency-first model architecture.
Source: Kog homepage.
https://www.kog.ai
-
2026-08-01:
Kog posts active hiring for GPU Engineer and Research Engineer
(hr)
. Kog said it is growing the team and shared active openings for a GPU Engineer and a Research Engineer, with work focused on CUDA/PTX, HIP/CDNA ISA, model morphing, and frontier MoE ports.
Source: Company social / careers evidence.
https://www.linkedin.com/posts/nicolas-constant1_were-growing-the-team-at-kog-and-im-hiring-activity-7478368546985709568-wDsr
-
2026-06-01:
Kog releases Laneformer 2B model and code
(product)
. Kog released the weights, model code, and documentation for Laneformer 2B on Hugging Face, describing it as a 2.3B-parameter instruction-tuned coding model designed for high-speed decoding.
Source: Official technical blog / model release.
https://blog.kog.ai/kog-laneformer-2b-the-latency-first-model-behind-kog-inference-engine
-
2026-05-28:
Kog launches tech preview of its inference engine
(product)
. Kog announced a public tech preview of the Kog Inference Engine with 3,000 output tokens/s per request on 8× AMD MI300X GPUs and 2,100 on 8× NVIDIA H200 GPUs, and made a playground available for testing.
Source: Official blog / company announcement.
https://blog.kog.ai/real-time-llm-inference-on-standard-gpus-3-000-tokens-s-per-request
-
2026-05-28:
Seed
(funding)
. Amount: USD 5M. Investors: Varsity VC, Bpifrance Deep Tech Program
Source: Official blog / company announcement.
https://blog.kog.ai/real-time-llm-inference-on-standard-gpus-3-000-tokens-s-per-request
-
2026-01-01:
AMD Developer highlights Kog collaboration
(gtm)
. AMD Developer shared a post featuring Gaël Delalleau, identified as co-founder and CEO of Kog, discussing how Kog optimizes AI inference on AMD GPUs.
Source: External partner/social post.
https://www.linkedin.com/posts/amd-developer_hear-from-ga%C3%ABl-delalleau-co-founder-and-activity-7474463730886492160-kAsl
Callosum
Website: callosum.com
Callosum builds software infrastructure for heterogeneous AI compute, with a focus on orchestrating inference across different chips and hardware stacks. ([callosum.com](https://www.callosum.com/?utm_source=openai))
Category: Inference engines, compilers & runtimes.
Funding stage: Seed.
Founded: 2025.
Country: United Kingdom.
Employees: 20-30.
Landscape fit: It fits this category because its core product is the runtime/orchestration layer that helps models run efficiently across heterogeneous hardware. The company says it is building inference systems and resource orchestration software, and its latest disclosed round was a $10.25M pre-seed led by Plural. ([callosum.com](https://www.callosum.com/join-us?utm_source=openai))
Funding rounds
-
seed round
announced 2026-08-20
: 100000000 USD
; investors: Atomico, Plural, DCVC, Sovereign AI, angel investors
; source: tech.eu
; https://tech.eu/2026/08/20/callosum-raises-100m-seed-round/
-
government investment
announced 2026-04-01
; investors: UK Sovereign AI Fund
; source: Callosum
; https://www.callosum.com/blog/sov-ai-investment
Company timeline
-
2026-08-20:
Callosum raises $100M seed round
(funding)
. Callosum raised $100 million in a seed round led by Atomico, with participation from Plural, DCVC, Sovereign AI, and angel investors.
Source: tech.eu.
https://tech.eu/2026/08/20/callosum-raises-100m-seed-round/
-
2026-08-20:
Product snapshot
(product)
. Callosum provides software infrastructure for heterogeneous AI compute, centered on orchestrating inference and workloads across different chips, models, and cloud environments.
Source: Callosum homepage.
https://www.callosum.com
-
2026-08-01:
Callosum expanded its compute partner network
(gtm)
. Callosum said it was partnering with next-generation silicon companies and infrastructure providers worldwide.
Source: Callosum.
https://www.callosum.com/blog/seed-round
-
2026-08-01:
Callosum pushed tailored inference into production
(product)
. Callosum published production results showing tailored inference applied with HelmGuard, including lower cost and faster sensitive-data detection.
Source: Callosum.
https://www.callosum.com/blog/tailored-inference
-
2026-08-01:
Callosum announced a flagship Cerebras partnership
(gtm)
. Callosum said it was working with Cerebras to bring low-latency heterogeneous multi-agent inference to customers at scale.
Source: Callosum.
https://www.callosum.com/blog/seed-round
-
2026-08-01:
Callosum made heterogeneity programmable
(product)
. Callosum described a compiler-like system that breaks workloads into blocks and routes each part to the model and chip that fits best.
Source: Callosum.
https://www.callosum.com/blog/programmable-heterogeneity
Wafer
Website: wafer.ai
Wafer builds AI agents and an inference platform that continuously optimize GPU serving stacks for faster, more efficient model inference.
Category: Inference engines, compilers & runtimes.
Funding stage: Series A.
Founded: 2025.
Country: United States.
Employees: 2-10 employees.
Landscape fit: It fits this category because it improves how trained models are compiled, scheduled, and served on underlying hardware rather than building the models themselves.
Funding rounds
-
Series A
announced 2026-09-01
: 40000000 USD
; investors: Marathon, Chemistry, Wing, AMD Ventures, Outset Capital, Fifty Years, Y Combinator, Jeff Dean, Guillermo Rauch, Andy Fang, Kyle Vogt, Akshay Kothari, Matthew Prince, Scott Stephenson
; source: Wafer blog
; https://www.wafer.ai/blog/series-a
-
seed
announced 2026-04-14
: 4000000 USD
; investors: Fifty Years, Liquid2, Y Combinator, Jeff Dean, Wojciech Zaremba, Arash Ferdowsi, Dan Fu
; source: Official Wafer blog
; https://www.wafer.ai/blog/seed-round
Company timeline
-
2026-09-03:
Wafer reports 2T+ tokens of continual inference
(milestone)
. Wafer's current homepage reports more than 2 trillion tokens of continual inference and presents production references including Vercel, AWS, Vapi, Sarvam AI, Tavus, Inworld, Brilliant, and DigitalOcean.
Source: Wafer homepage.
https://www.wafer.ai/
-
2026-09-03:
HR snapshot
(hr)
. Wafer appears to be a small company with active hiring. The public careers signal points to open roles across engineering, go to market, and operations.
Source: Y Combinator jobs page.
https://www.ycombinator.com/companies/wafer/jobs
-
2026-09-03:
Product snapshot
(product)
. Wafer builds AI agents that optimize GPU kernels and an inference platform for faster model serving. The current surface includes serverless and dedicated inference for open source LLMs, plus IDE and CLI tools for profiling, compiler inspection, GPU documentation search, trace analysis, and kernel benchmarking.
Source: Wafer homepage.
https://www.wafer.ai
-
2026-09-01:
Wafer receives acquisition offers as valuation tops $200M
(milestone)
. The Information reported that Wafer had received acquisition offers and that its Series A valued the company above $200 million.
Source: The Information.
https://www.theinformation.com/newsletters/ai-agenda/wafer-inference-provider-uses-non-nvidia-chips-lands-acquisition-offers-200-million-plus-valuation
-
2026-09-01:
Series A
(funding)
. Amount: USD 40M. Investors: Marathon, Chemistry, Wing, AMD Ventures, Outset Capital, Fifty Years, Y Combinator, Jeff Dean, Guillermo Rauch, Andy Fang, Kyle Vogt, Akshay Kothari, Matthew Prince, Scott Stephenson
Source: Wafer blog.
https://www.wafer.ai/blog/series-a
-
2026-07-17:
Wafer integration with TrueFoundry
(gtm)
. Wafer announced an integration with TrueFoundry AI Gateway for unified routing, observability, and zero data retention.
Source: Wafer blog.
https://www.wafer.ai/blog
Model serving & deployment platforms
Platforms that help teams deploy and operate models as reliable production services while handling scaling and infrastructure complexity.
Baseten
Website: baseten.co
Baseten is an AI inference infrastructure platform that helps teams train, deploy, and serve models in production with autoscaling, observability, and optimized serving infrastructure.
Category: Model serving & deployment platforms.
Funding stage: Series D+.
Founded: 2019.
Country: United States.
Employees: 201-500 employees.
Landscape fit: It fits the model serving & deployment category because its core product is an inference platform for turning models into production APIs and operating them reliably at scale, including autoscaling, observability, and multi-cloud/self-hosted deployment support.
Funding rounds
-
Series F
announced 2026-06-01
: 1500000000 USD
; investors: Altimeter Capital, Conviction Partners, Spark Capital, Sands Capital, Wellington Management, Battery Ventures, Blackbird, D.E. Shaw Ventures, Durable Capital Partners, Greylock, IVP, Verified Capital, 01A
; source: Baseten
; https://www.baseten.co/blog/announcing-our-series-f/
-
Series E
announced 2026-01-01
: 300000000 USD
; investors: IVP, CapitalG, 01A, Altimeter, Battery Ventures, BOND, BoxGroup, Blackbird Ventures, Conviction, Greylock, NVIDIA
; source: Baseten
; https://www.baseten.co/blog/announcing-baseten-s-300m-series-e/
-
Series D
announced 2025-09-01
: 150000000 USD
; investors: BOND, Conviction, CapitalG, 01A, IVP, Spark, Greylock, Scribble Ventures, BoxGroup, Premji Invest
; source: Baseten
; https://www.baseten.co/blog/announcing-baseten-150m-series-d/
-
Series C
announced 2025-02-01
: 75000000 USD
; investors: IVP, Spark Capital, Greylock, Conviction, South Park Commons, Basecase, Lachy Groom, 01A
; source: Baseten
; https://www.baseten.co/blog/announcing-baseten-75m-series-c/
-
Series B
announced 2024-03-01
: 40000000 USD
; investors: IVP, Spark Capital, Greylock, South Park Commons, Lachy Groom, Base Case, Conviction
; source: Baseten
; https://www.baseten.co/blog/announcing-our-series-b/
Company timeline
-
2026-07-01:
Broad open-role hiring cluster visible on LinkedIn
(hr)
. Baseten's jobs page showed multiple current openings across engineering, sales, field productivity and enablement, product marketing, and strategic sales roles.. Team size observed: 201-500 employees employees.
Source: LinkedIn Jobs.
https://www.linkedin.com/company/baseten/jobs
-
2026-06-11:
Mercury 2 launch reinforces Baseten’s model-labs and production proof motion
(gtm)
. Baseten announced Mercury 2 on Baseten and highlighted production usage, including a white-labeled deployment path and performance results for customer traffic.
Source: Baseten blog.
https://www.baseten.co/blog/mercury-2-is-now-available-on-baseten/
-
2026-06-03:
Gabe Stern joins as General Counsel
(hr)
. Baseten announced Gabe Stern as General Counsel to build the company's legal foundation.. Team size observed: 201-500 employees employees.
Source: Baseten blog.
https://www.baseten.co/blog/welcome-gabe-stern/
-
2026-06-01:
Series F
(funding)
. Amount: 1500000000 USD. Investors: Altimeter Capital, Conviction Partners, Spark Capital, Sands Capital, Wellington Management, Battery Ventures, Blackbird, D.E. Shaw Ventures, Durable Capital Partners, Greylock, IVP, Verified Capital, 01A
Source: Baseten.
https://www.baseten.co/blog/announcing-our-series-f/
-
2026-05-06:
Frontier Gateway launch
(product)
. Baseten launched Frontier Gateway, a managed routing layer on top of Dedicated Inference for production-grade multi-tenant inference APIs under a customer domain.
Source: Baseten blog.
https://www.baseten.co/blog/introducing-baseten-frontier-gateway/
-
2026-05-06:
Frontier Gateway targets model labs with white-labeled multi-tenant APIs
(gtm)
. Baseten introduced Frontier Gateway, a managed routing layer for model labs to serve hosted models under their own domain without building a separate gateway.
Source: Baseten blog.
https://www.baseten.co/blog/introducing-baseten-frontier-gateway/
Fireworks AI
Website: fireworks.ai
Fireworks AI is an AI inference cloud and developer platform that helps teams deploy, serve, fine-tune, and optimize generative models in production with an emphasis on speed, cost efficiency, and control.
Category: Model serving & deployment platforms.
Funding stage: Series C.
Founded: 2022.
Country: United States.
Employees: 265 employees visible on LinkedIn.
Landscape fit: It fits model serving & deployment platforms because its core product is production inference and deployment infrastructure for running open and custom models reliably at scale, including serverless and dedicated deployments. ([fireworks.ai](https://fireworks.ai/?utm_source=openai))
Funding rounds
-
Series C
announced 2025-10-01
: 250000000 USD
; investors: Lightspeed Venture Partners, Index Ventures, Evantic, Sequoia Capital
; source: Fireworks AI Blog
; https://fireworks.ai/blog/series-c
-
Series B
announced 2024-07-01
: 52000000 USD
; investors: Sequoia Capital, NVIDIA, AMD, MongoDB Ventures, Benchmark
; source: Fireworks AI Blog
; https://fireworks.ai/blog/fireworks-ai-series-b-compound-ai
-
Series A
announced 2024-03-01
: 25000000 USD
; investors: Benchmark, Sequoia Capital, Databricks Ventures, Frank Slootman, Alexandr Wang, Sheryl Sandberg, Howie Liu, Artisanal Ventures
; source: FinSMEs
; https://www.finsmes.com/2024/03/fireworks-ai-raises-25m-in-series-a-funding.html
Company timeline
-
2026-07-01:
Careers-page hiring snapshot
(hr)
. Fireworks AI’s careers page currently shows 36 open positions.. Team size observed: 265 employees visible on LinkedIn employees.
Source: Fireworks AI Careers.
https://fireworks.ai/careers
-
2026-07-01:
Hiring signal: AI field engineer role
(hr)
. A candidate publicly announced joining Fireworks AI as an AI Field Engineer.. Team size observed: 265 employees visible on LinkedIn employees.
Source: LinkedIn.
https://www.linkedin.com/posts/peter-t-300b9760_excited-to-share-that-ill-be-joining-fireworks-activity-7477482612257619968-XAo-
-
2026-07-01:
Hiring cluster: partner-team buildout
(hr)
. Peter Kuo said Fireworks AI had new partner-team roles open, including Head of Systems Integrators, AI Field Engineers, Strategic Alliances, Partner Sales Manager (Microsoft), and Partner Development Manager (AWS).. Team size observed: 265 employees visible on LinkedIn employees.
Source: LinkedIn.
https://www.linkedin.com/posts/peterzkuo_one-emerging-maxim-in-ai-is-leads-are-ephemeral-activity-7476273138641592320-n3Kw
-
2026-06-01:
Trilogy customer story reinforces enterprise production-scale adoption
(gtm)
. Fireworks published a customer story on Trilogy, describing how open-weight model usage was unified across internal teams and portfolio workflows for production-scale agentic systems.
Source: Fireworks AI blog.
https://fireworks.ai/blog/Trilogy
-
2026-05-26:
Serverless 2.0 introduces tiered serving paths
(product)
. Fireworks launched Serverless 2.0 with Standard, Priority, and Fast serving paths on one API surface, plus clearer error semantics and a preview of Background async processing.
Source: Fireworks AI Blog.
https://fireworks.ai/blog/serverless-2
-
2026-05-01:
Hiring signal: technical developer advocate opening
(hr)
. Aishwarya Srinivasan posted that Fireworks AI was hiring a Technical Developer Advocate.. Team size observed: 265 employees visible on LinkedIn employees.
Source: LinkedIn.
https://www.linkedin.com/posts/aishwarya-srinivasan_were-hiring-a-technical-developer-advocate-activity-7346341920920514560-WCK8
BentoML
Website: bentoml.com
BentoML is a unified inference platform and open-source framework for deploying, scaling, and operating AI models in production with infrastructure, autoscaling, observability, and cloud/on-prem deployment options. ([docs.bentoml.com](https://docs.bentoml.com/en/latest/?utm_source=openai))
Category: Model serving & deployment platforms.
Funding stage: Acquired.
Founded: 2019.
Country: United States.
Employees: 51-200 employees.
Landscape fit: It fits model serving & deployment because BentoML explicitly focuses on deploying and scaling AI models as production services, with inference, autoscaling, monitoring, and multi-cloud/on-prem deployment features. ([docs.bentoml.com](https://docs.bentoml.com/en/latest/?utm_source=openai))
Funding rounds
-
seed
announced 2023-06-01
: 9000000 USD
; investors: DCM Ventures, Bow Capital, Firestreak Ventures
; source: TechCrunch
; https://techcrunch.com/2023/06/26/bentoml-scores-9m-funding-to-expedite-ai-app-development/
Company timeline
-
2026-02-10:
BentoML joined Modular
(gtm)
. BentoML announced it had joined Modular after months of collaborating on customer deployments, describing the move as a strategic product acquisition and promising continued support for the open-source project and enterprise inference customers. ([bentoml.com](https://www.bentoml.com/blog/bentoml-is-joining-modular))
Source: BentoML blog.
https://www.bentoml.com/blog/bentoml-is-joining-modular
-
2026-02-10:
BentoML joins Modular
(product)
. BentoML announced it was joining Modular to build the next generation of AI inference infrastructure while continuing support for multi-cloud, BYOC, and the open-source project.
Source: BentoML blog.
https://www.bentoml.com/blog/bentoml-is-joining-modular
-
2026-02-10:
BentoML announced it was joining Modular
(hr)
. Chaoyu Yang and the BentoML team announced that BentoML had joined Modular as part of a strategic product acquisition.. Team size observed: 51-200 employees employees.
Source: BentoML blog.
https://www.bentoml.com/blog/bentoml-is-joining-modular
-
2026-02-01:
BentoML acquired by Modular
(acquisition)
. BentoML announced it was joining Modular as part of a strategic product acquisition.. Acquirer: Modular. Status: announced
Source: BentoML blog.
https://www.bentoml.com/blog/bentoml-is-joining-modular
-
2025-12-28:
BentoML shifts LLM serving toward OpenAI-compatible vLLM workflows
(product)
. BentoML’s vLLM docs and a 2025 tutorial showed OpenAI-compatible endpoints, deployment to BentoCloud, autoscaling, scale-to-zero, and observability for private reasoning APIs.
Source: BentoML blog.
https://www.bentoml.com/blog/deploying-a-large-language-model-with-bentoml-and-vllm
-
2025-09-17:
BYOC and compliance-first enterprise messaging expanded
(gtm)
. BentoML followed with a BYOC-and-on-prem article that argued enterprises can combine managed-service speed with data ownership, compliance, and private-cloud control, reinforcing a clear enterprise sales narrative around security-sensitive inference. ([bentoml.com](https://www.bentoml.com/blog/how-enterprises-can-scale-ai-securely-with-byoc-and-on-prem-deployments))
Source: BentoML blog and docs.
https://www.bentoml.com/blog/how-enterprises-can-scale-ai-securely-with-byoc-and-on-prem-deployments
FriendliAI
Website: friendli.ai
FriendliAI is a generative AI inference platform that helps teams deploy and operate large language and multimodal models in production with optimized speed, latency, and cost.
Category: Model serving & deployment platforms.
Funding stage: Seed.
Founded: 2021.
Country: South Korea.
Employees: 11-50 employees.
Landscape fit: It fits the model serving & deployment platforms category because its core product is production inference infrastructure for serving models reliably at scale, including APIs, dedicated endpoints, and private container deployments.
Funding rounds
-
Seed extension
announced 2025-08-01
: 20000000 USD
; investors: Capstone Partners, Sierra Ventures, Alumni Ventures, KDB, KB Securities
; source: FriendliAI
; https://friendli.ai/blog/friendliai-raises-20m-in-seed-extension-round
Company timeline
-
2026-07-01:
Kilo Code partnership for coding agents
(gtm)
. FriendliAI partnered with Kilo Code and NVIDIA Nemotron to bring open-source coding agents into production through FriendliAI’s inference layer.
Source: FriendliAI Blog.
https://friendli.ai/blog/Kilo-Code-FriendliAI-NVIDIA-Nemotron
-
2026-07-01:
Current careers page shows 18 open roles
(hr)
. FriendliAI's careers page currently lists 18 open positions spanning engineering, infrastructure, inference systems, solutions, marketing, product, and sales across Seoul and San Francisco.. Team size observed: 11-50 employees employees.
Source: FriendliAI Careers Page.
https://friendli.ai/careers
-
2026-06-01:
Hana Rhee joined as VP of Sales
(hr)
. Hana Rhee publicly said she joined FriendliAI as Vice President of Sales.. Team size observed: 11-50 employees employees.
Source: LinkedIn.
https://www.linkedin.com/posts/hanarhee_excited-to-share-that-ive-joined-friendliai-activity-7462153648740069376-Lf5j
-
2026-05-11:
San Francisco office expansion
(gtm)
. FriendliAI opened a San Francisco office to deepen proximity to customers and developers in the AI ecosystem.
Source: FriendliAI Blog.
https://friendli.ai/blog/friendliai-sf-office
-
2026-04-15:
Anthropic Messages API compatibility
(gtm)
. FriendliAI added Anthropic Messages API support across Serverless and Dedicated Endpoints and offered migration credits for teams moving from other providers.
Source: FriendliAI Blog.
https://friendli.ai/blog/friendliai-supports-anthropic-messages-api
-
2026-04-15:
Anthropic Messages API support added
(product)
. FriendliAI added Anthropic Messages API compatibility across both Serverless and Dedicated Endpoints, enabling Claude-based applications to switch base URL and token while using open-weight models. ([friendli.ai](https://friendli.ai/blog/friendliai-supports-anthropic-messages-api?utm_source=openai))
Source: FriendliAI Blog.
https://friendli.ai/blog/friendliai-supports-anthropic-messages-api
Modal
Website: modal.com
Modal is a serverless cloud platform for AI/ML workloads that helps teams deploy and operate code, GPU jobs, inference, fine-tuning, and other production compute without managing infrastructure. ([modal.com](https://modal.com/company?utm_source=openai))
Category: Model serving & deployment platforms.
Funding stage: Series C.
Founded: 2021.
Country: United States.
Employees: About 192 employees.
Landscape fit: It fits model serving & deployment platforms because Modal explicitly supports low-latency elastic inference, GPU-enabled containers, autoscaling production workloads, and AI inference/fine-tuning use cases, which are core model-serving/deployment capabilities. ([modal.com](https://modal.com/blog/modal-series-c?utm_source=openai))
Funding rounds
-
Series C
announced 2026-05-01
: 355000000 USD
; investors: General Catalyst, Redpoint Ventures, Menlo Ventures, Bain Capital Ventures, Accel, existing major investors
; source: Modal
; https://modal.com/blog/modal-series-c
-
Series B
announced 2025-09-01
: 87000000 USD
; investors: Lux Capital, Amplify Partners, Redpoint Ventures
; source: Modal
; https://modal.com/blog/announcing-our-series-b
-
Series A
announced 2023-10-01
: 16000000 USD
; investors: Redpoint Ventures, Amplify Partners, Lux Capital, Definition Capital
; source: Modal
; https://modal.com/blog/general-availability-and-series-a-press-release
-
Seed
announced 2022-02-01
: 7000000 USD
; investors: Amplify Partners
; source: Amplify Partners
; https://www.amplifypartners.com/blog-posts/modal
Company timeline
-
2026-07-06:
Serverless GPU pricing education content
(gtm)
. Modal published an interactive guide on pricing serverless GPUs for inference, training, and agentic development, positioning its economics against reserved capacity. ([modal.com](https://modal.com/blog/how-to-price-serverless?utm_source=openai))
Source: Modal blog.
https://modal.com/blog/how-to-price-serverless
-
2026-06-30:
Anthropic Claude Science integration
(product)
. Modal announced an integration with Anthropic’s Claude Science for compute-heavy life sciences workflows.
Source: Modal Blog.
https://modal.com/blog/modal-integration-brings-scalable-compute-to-claude-science
-
2026-06-23:
Modal Auto Endpoints
(product)
. Modal launched OpenAI API-compatible Auto Endpoints for production inference.
Source: Modal Blog.
https://modal.com/blog/introducing-auto-endpoints
-
2026-05-27:
RBAC for humans and agents
(product)
. Modal launched role-based access control with restricted environments for Team and Enterprise plans.
Source: Modal Blog.
https://modal.com/blog/role-based-access-control-for-humans-and-agents
-
2026-05-01:
Series C
(funding)
. Amount: USD 355M. Investors: General Catalyst, Redpoint Ventures, Menlo Ventures, Bain Capital Ventures, Accel, existing major investors
Source: Modal.
https://modal.com/blog/modal-series-c
-
2026-04-10:
Butter founder and researcher join Modal Sandbox team
(hr)
. Modal announced that Butter founder Erik Dunteman and researcher Raymond Tana would join the Modal Sandbox team as part of the acquisition.. Team size observed: About 192 employees employees.
Source: Modal blog.
https://modal.com/blog/butter-is-joining-modal
Replicate
Website: replicate.com
A platform that lets developers run, fine-tune, and deploy machine-learning models through an API without managing underlying infrastructure.
Category: Model serving & deployment platforms.
Funding stage: Acquired.
Founded: 2019.
Country: United States.
Employees: 26 employees.
Landscape fit: It fits model serving & deployment because Replicate provides cloud hosting, API access, scaling, and deployment tooling for production ML models, handling the infrastructure complexity for teams that want to ship reliable inference services. ([replicate.com](https://replicate.com/blog/hello-world/?utm_source=openai))
Funding rounds
-
Series B
announced 2023-12-01
: 40000000 USD
; investors: Andreessen Horowitz, NVentures, Heavybit, Sequoia Capital, Y Combinator
; source: Replicate Blog
; https://replicate.com/blog/series-b
-
Series A
announced 2023-02-01
: 12500000 USD
; investors: Andreessen Horowitz, Y Combinator, Sequoia Capital, Dylan Field, Guillermo Rauch
; source: Andreessen Horowitz
; https://a16z.com/announcement/investing-in-replicate/
-
Seed
announced 2023-02-01
: 5300000 USD
; investors: Sequoia Capital
; source: Sequoia Capital
; https://www.sequoiacap.com/article/partnering-with-replicate-machine-learning-simplified/
Company timeline
-
2026-04-21:
Agent skills for Replicate
(product)
. Replicate published agent skills that teach coding assistants how to discover, compare, and use models on Replicate. ([replicate.com](https://replicate.com/changelog))
Source: Replicate changelog.
https://replicate.com/changelog
-
2026-04-21:
Agent skills for Replicate
(gtm)
. Replicate published agent skills that teach coding assistants how to discover, compare, and use models on the platform. ([replicate.com](https://replicate.com/changelog))
Source: Replicate changelog.
https://replicate.com/changelog
-
2026-02-10:
MCP server auto-discovery
(product)
. Replicate added `/.well-known/mcp/server.json` so its MCP server can be discovered through the official MCP Registry. ([replicate.com](https://replicate.com/changelog))
Source: Replicate changelog.
https://replicate.com/changelog
-
2026-02-10:
MCP server auto-discovery via the MCP Registry
(gtm)
. Replicate made its MCP server discoverable through the official MCP Registry using a server.json endpoint. ([replicate.com](https://replicate.com/changelog?utm_source=openai))
Source: Replicate changelog.
https://replicate.com/changelog
-
2026-01-14:
Prediction list filter by source
(product)
. Replicate added a `source=web` filter to the predictions API so users can isolate runs created in the web interface. ([replicate.com](https://replicate.com/changelog))
Source: Replicate changelog.
https://replicate.com/changelog
-
2025-11-17:
Replicate joined Cloudflare
(gtm)
. Replicate announced it was joining Cloudflare while remaining a distinct brand and integrating with Cloudflare’s Developer Platform. ([replicate.com](https://replicate.com/blog/replicate-cloudflare/))
Source: Replicate blog.
https://replicate.com/blog/replicate-cloudflare/
Together AI
Website: together.ai
Together AI is a full-stack AI cloud platform that helps teams run, fine-tune, train, and deploy open-source models and AI agents with production-grade inference and infrastructure.
Category: Model serving & deployment platforms.
Funding stage: Series C.
Founded: 2022.
Country: United States.
Employees: 201-500 employees.
Landscape fit: It fits model serving & deployment platforms because its core offering is production inference and deployment infrastructure for open-source AI models, with scaling, optimization, and managed compute handled by the platform.
Funding rounds
-
Series C
announced 2026-07-01
: 800000000 USD
; investors: Aramco Ventures, NVIDIA, Vista Equity, General Catalyst, Emergence Capital, SE Ventures, Pegatron, Salesforce Ventures, March Capital, DTCP Growth, Lux Capital, Geodesic, PSP Partners
; source: Together AI
; https://www.together.ai/blog/announcing-our-series-c
-
Series B
announced 2025-02-01
: 305000000 USD
; investors: General Catalyst, Prosperity7, Salesforce Ventures, DAMAC Capital, NVIDIA, Kleiner Perkins, March Capital, Emergence Capital, Lux Capital, SE Ventures, Greycroft, Coatue, Definition, Cadenza Ventures, Long Journey Ventures, Brave Capital, Scott Banister, SK Telecom, John Chambers
; source: Together AI
; https://www.together.ai/blog/together-ai-announcing-305m-series-b
-
Series A
announced 2024-03-01
: 106000000 USD
; investors: Salesforce Ventures, Coatue, Lux Capital, Kleiner Perkins, Emergence Capital, Prosperity7 Ventures, NEA, Greycroft, Definition Capital, Long Journey Ventures, Factory, Scott Banister, SV Angel, Clem Delangue, Soumith Chintala
; source: Together AI
; https://www.together.ai/blog/series-a2
-
Series A extension
announced 2024-03-01
: 106000000 USD
; investors: Salesforce Ventures, Coatue, Lux Capital, Kleiner Perkins, Emergence Capital, Prosperity7 Ventures, NEA, Greycroft, Definition Capital, Long Journey Ventures, Factory, SV Angel
; source: Together AI blog
; https://www.together.ai/blog/series-a2
-
Series A
announced 2023-11-01
: 102500000 USD
; investors: Kleiner Perkins, NVIDIA, Emergence Capital, NEA, Prosperity7 Ventures, Greycroft
; source: TechCrunch
; https://techcrunch.com/2023/11/29/together-lands-102-5m-investment-to-grow-its-cloud-for-training-generative-ai/
Company timeline
-
2026-07-10:
Fresh multi-role hiring cluster appears on LinkedIn
(hr)
. Together AI’s LinkedIn jobs page shows multiple roles posted within the last 10 to 24 hours, including customer success, customer support, site reliability leadership, data center operations, analytics engineering, and research internships.. Team size observed: 201-500 employees employees.
Source: LinkedIn.
https://www.linkedin.com/company/togethercomputer/jobs
-
2026-07-08:
Provisioned Throughput launched for reserved inference capacity
(product)
. Together AI introduced Provisioned Throughput, a new inference form factor for frontier open models that combines token-based pricing with reserved capacity and a 99% uptime SLA.
Source: Together AI Blog.
https://www.together.ai/blog/provisioned-throughput
-
2026-07-07:
ICML 2026 post explicitly recruits researchers and research engineers
(hr)
. Together AI’s ICML 2026 blog post says the company is hiring researchers and research engineers and directs candidates to open roles on its careers page.. Team size observed: 201-500 employees employees.
Source: Together AI.
https://www.together.ai/blog/icml-2026
-
2026-07-01:
Series C
(funding)
. Amount: 800000000 USD. Investors: Aramco Ventures, NVIDIA, Vista Equity, General Catalyst, Emergence Capital, SE Ventures, Pegatron, Salesforce Ventures, March Capital, DTCP Growth, Lux Capital, Geodesic, PSP Partners
Source: Together AI.
https://www.together.ai/blog/announcing-our-series-c
-
2026-05-15:
Together AI launches Pearl Research Labs inference partnership
(gtm)
. Together AI announced an exclusive partnership with Pearl Research Labs and a discounted Gemma-4-31B-it-pearl inference endpoint tied to Pearl’s token economics. ([together.ai](https://www.together.ai/blog/together-ai-partners-with-pearl-research-labs?utm_source=openai))
Source: Together AI Blog.
https://www.together.ai/blog/together-ai-partners-with-pearl-research-labs
-
2026-04-30:
Adaption partnership exposed Together fine-tuning through Adaptive Data
(product)
. Together partnered with Adaption so users could connect Together Fine-Tuning inside Adaption’s workflow.
Source: Together AI Blog.
https://www.together.ai/blog/announcing-together-ai-and-adaption-partnership
Runpod
Website: runpod.io
Runpod is an AI developer cloud that lets teams build, train, fine-tune, deploy, and scale AI workloads on GPU infrastructure from one platform.
Category: Model serving & deployment platforms.
Funding stage: Series A.
Founded: 2022.
Country: United States.
Employees: 107 employees.
Landscape fit: It fits the model serving & deployment platforms category because its core offering helps users deploy and operate AI models in production, including serverless inference, scalable GPU compute, and deployment tooling.
Funding rounds
-
Series A
announced 2026-06-24
: 100000000 USD
; investors: Summit Partners
; source: PR Newswire
; https://www.prnewswire.com/news-releases/runpod-raises-100m-led-by-summit-partners-to-accelerate-the-ai-developer-cloud-302808689.html
-
Seed
announced 2024-05-01
: 20000000 USD
; investors: Intel Capital, Dell Technologies Capital, Julien Chaummond, Nat Friedman, Adam Lewis
; source: Business Wire
; https://www.businesswire.com/news/home/20240508053225/en/RunPod-Raises-%2420M-in-Seed-Funding-Co-led-by-Intel-Capital-and-Dell-Technologies-Capital
Company timeline
-
2026-07-10:
Hiring cluster expands across engineering, infrastructure, marketing, people, product, sales, and support
(hr)
. Runpod’s careers pages show a broad set of current openings, including Director of Infrastructure Engineering, Director of Software Engineering - Product & Platform Delivery, Senior Product Marketing Manager, Head of People Operations, Senior Product Manager, and multiple sales and infrastructure roles.. Team size observed: 107 employees employees.
Source: LinkedIn Jobs / Greenhouse.
https://www.linkedin.com/company/runpod-io/jobs
-
2026-07-01:
Runpod Overdrive launched
(product)
. Runpod launched Overdrive, an inference optimization engine for production LLM inference on Serverless that is tuned to a specific model and workload.
Source: Runpod Blog.
https://www.runpod.io/blog/introducing-runpod-overdrive
-
2026-07-01:
Runpod launches Overdrive for optimized production inference
(gtm)
. Runpod introduced Overdrive, an inference optimization engine for teams running production LLM inference on Serverless, with sales-assisted access and throughput and latency gains used as the headline proof point.
Source: Runpod blog.
https://www.runpod.io/blog/introducing-runpod-overdrive
-
2026-06-25:
Runpod Serverless added faster cold starts, batch inference, and no-Docker deploys
(product)
. Runpod described major upgrades to Serverless, including FlashBoot sub-200ms cold starts, batch inference via run-batch, and a no-Docker Flash deployment path with model-first defaults.
Source: Runpod Blog.
https://www.runpod.io/blog/whats-new-in-runpod-serverless-faster-cold-starts-batch-inference-and-no-docker-deploys
-
2026-06-24:
Series A
(funding)
. Amount: 100000000 USD. Investors: Summit Partners
Source: PR Newswire.
https://www.prnewswire.com/news-releases/runpod-raises-100m-led-by-summit-partners-to-accelerate-the-ai-developer-cloud-302808689.html
-
2026-06-24:
AI developer cloud repositioning in Series A announcement
(gtm)
. Runpod used its Series A announcement to frame the company as an AI developer cloud spanning the full lifecycle: build, train, fine-tune, deploy, and scale, alongside proof that more than one million developers are building on the platform.
Source: Runpod blog.
https://www.runpod.io/blog/one-million-developers
OpenRelay
Website: openrelay.inc
OpenRelay is a distributed, hardware-agnostic AI inference platform that connects consumer and datacenter GPUs into a fault-tolerant mesh for deploying production model workloads.
Category: Model serving & deployment platforms.
Funding stage: Pre Seed.
Founded: 2026.
Country: United States.
Employees: 2 employees.
Landscape fit: It fits model serving & deployment platforms because it provides infrastructure to deploy, scale, and operate inference workloads in production, including routing, failover, and GPU orchestration.
Company timeline
-
2026-07-01:
Public careers presence without visible openings
(hr)
. OpenRelay’s site includes a Careers section inviting candidates to get in touch, but the checked official pages did not show any named open roles or a hiring list.. Team size observed: 2 employees employees.
Source: OpenRelay.
https://openrelay.inc/about
-
2026-06-10:
Pricing and availability logic tightened
(gtm)
. OpenRelay updated its availability API and dashboard so only satisfiable GPU counts can be selected and silent deploy timeouts are rejected upfront.
Source: OpenRelay Product Updates.
https://openrelay.inc/updates
-
2026-06-10:
Docs and API reference formalized
(gtm)
. OpenRelay launched a documentation site generated from the control-plane OpenAPI spec with 89 customer-facing operations and hand-written guides.
Source: OpenRelay Product Updates.
https://openrelay.inc/updates
-
2026-06-10:
GPU placement validation and availability granularity added
(product)
. OpenRelay updated the availability API to expose GPU allocation granularity, reject unsatisfiable GPU counts at create time, and constrain the dashboard wizard to valid counts.
Source: OpenRelay Product Updates.
https://openrelay.inc/updates
-
2026-06-10:
API docs site launched from the OpenAPI spec
(product)
. OpenRelay launched a new documentation site generated from the control-plane OpenAPI spec, with 89 customer-facing operations and hand-written guides for authentication, errors, pagination, and webhooks.
Source: OpenRelay Product Updates.
https://openrelay.inc/updates
-
2026-06-08:
Production launch with billing and docs
(gtm)
. OpenRelay said its production web tier, dashboard, control-plane API, and live Stripe billing were now live.
Source: OpenRelay Product Updates.
https://openrelay.inc/updates
Blackfuel
Website: blackfuel.ai
Blackfuel is an AI inference company that provides dedicated infrastructure and an OpenAI compatible API for serving open weight models in production.
Category: Model serving & deployment platforms.
Founded: 2026.
Country: France.
Employees: unknown.
Landscape fit: It fits model serving and deployment platforms because its core product is an inference API and dedicated production capacity for running models reliably at scale.
Company timeline
-
2026-10-01:
HR snapshot
(hr)
. The available evidence identifies a co-founder and includes a LinkedIn source summary associating Dali Kilani with Blackfuel. It does not establish a current headcount or hiring signal.
Source: Blackfuel launch announcement.
https://www.blackfuel.ai/press.html
-
2026-10-01:
Product snapshot
(product)
. Blackfuel documents an OpenAI-compatible inference API for open-weight models, together with an infrastructure offering described in its launch materials. The documentation covers chat completions, embeddings, model listing, streaming, reasoning controls, MCP-based account management, and security and privacy commitments.
Source: Blackfuel homepage.
https://www.blackfuel.ai
-
2026-09-30:
Blackfuel publishes reasoning controls
(product)
. Blackfuel documented reasoning content and a reasoning_effort parameter for supported models. The guide shows users can tune how the service handles reasoning output.
Source: Blackfuel reasoning guide.
https://docs.blackfuel.ai/docs/guides/reasoning
-
2026-09-30:
Blackfuel announces planned Barcelona deployment
(gtm)
. Blackfuel announced its first inference only deployment in Barcelona, Spain. The release names Digital Realty as host and says Dell and NTT DATA are supporting deployment and integration work.
Source: Blackfuel launch announcement.
https://www.blackfuel.ai/press.html
-
2026-09-30:
Blackfuel documents MCP account management
(product)
. Blackfuel published an MCP guide that lets compatible AI clients manage accounts through the service. The guide shows the product extending beyond basic inference into account operations.
Source: Blackfuel MCP guide.
https://docs.blackfuel.ai/docs/guides/mcp
-
2026-09-30:
Dali Kilani publicly mentions Blackfuel
(hr)
. A LinkedIn source summary says Dali Kilani described Blackfuel as a new French startup focused on making AI useful at scale through infrastructure and software. The fetched post text was unavailable, so the event is supported only by the summary record.
Source: Dali Kilani LinkedIn post.
https://de.linkedin.com/in/dalikilani
Inference gateways, routing & control
Infrastructure that connects applications with model providers and provides a shared control layer for production inference.
OpenRouter
Website: openrouter.ai
OpenRouter is an AI gateway and model marketplace that lets developers and enterprises route requests across hundreds of LLMs through a single API with built-in reliability, pricing, and fallback controls.
Category: Inference gateways, routing & control.
Funding stage: Acquired.
Founded: 2023.
Country: United States.
Employees: 11-50.
Landscape fit: It fits the inference gateways, routing & control category because it provides a shared control layer and unified API for choosing, routing, and managing production model inference across providers.
Funding rounds
-
Series B
announced 2026-05-26
: 113000000 USD
; investors: CapitalG, NVentures, ServiceNow Ventures, MongoDB Ventures, Snowflake Ventures, Databricks Ventures, AMP PBC, Pace Capital, Andreessen Horowitz, Menlo Ventures
; source: Business Wire
; https://www.businesswire.com/news/home/20260526953416/en/OpenRouter-Raises-%24113-Million-CapitalG-led-Series-B-as-Weekly-Volume-Explodes-to-25T-Tokens
-
Seed and Series A
announced 2025-06-25
: 40000000 USD
; investors: Andreessen Horowitz, Menlo Ventures, Sequoia, prominent industry angels
; source: OpenRouter raises $40 million to scale up multi-model inference for enterprise
; https://www.globenewswire.com/news-release/2025/06/25/3105125/0/en/openrouter-raises-40-million-to-scale-up-multi-model-inference-for-enterprise.html
-
Series Seed
announced 2024-12-01
: 10750000 USD
; source: Forge
; https://forgeglobal.com/openrouter_stock/
Company timeline
-
2026-09-07:
HR snapshot
(hr)
. OpenRouter shows an active multi function hiring posture, with openings spanning engineering, go to market, legal, marketing, operations, and product. The public hiring mix suggests a small but scaling team.
Source: Careers at OpenRouter.
https://openrouter.ai/careers?ashby_jid=7c5cb1ca-71ba-464d-bf98-33743e8b1474
-
2026-09-07:
Product snapshot
(product)
. OpenRouter provides a single API and gateway for accessing and routing across hundreds of AI models. The platform includes fallback handling, provider selection, pricing and cost controls, observability, workspaces, guardrails, response caching, analytics, evaluation tools, model benchmarks, and support for text, image, audio, speech, transcription, embeddings, video, search, and agent workflows.
Source: OpenRouter Docs.
https://openrouter.ai/docs/quickstart
-
2026-08-19:
Stripe announces acquisition of OpenRouter
(acquisition)
. Stripe announced that it has agreed to acquire OpenRouter, the AI model gateway and routing platform. The transaction was publicly announced by Stripe on August 19, 2026.
Source: Stripe.
https://stripe.com/newsroom/news/stripe-agrees-to-acquire-openrouter
-
2026-06-25:
OpenRouter MCP Server launch
(gtm)
. OpenRouter released an MCP server that exposes live model catalog data, benchmarks, pricing, docs search, and test inference directly inside coding agents and editors.
Source: OpenRouter Blog.
https://openrouter.ai/blog/announcements/openrouter-mcp-server/
-
2026-06-25:
OpenRouter MCP server launched
(product)
. OpenRouter released an MCP server that exposes live model catalog data, benchmark rankings, pricing, docs, and test inference directly inside coding agents and other MCP clients.
Source: OpenRouter Blog.
https://openrouter.ai/blog/announcements/openrouter-mcp-server/
-
2026-06-23:
Unified Image API launch
(gtm)
. OpenRouter launched a dedicated image-generation API with unified access to 30+ models across multiple providers, standardized request shapes, and capability discovery for model-specific parameters.
Source: OpenRouter Blog.
https://openrouter.ai/blog/announcements/image-api/
Not Diamond
Website: notdiamond.ai
Not Diamond builds an intelligent AI model router and prompt optimization platform that helps developers route requests across multiple models to improve quality while reducing inference cost and latency.
Category: Inference gateways, routing & control.
Funding stage: Seed.
Founded: 2023.
Country: United States.
Employees: 11-50.
Landscape fit: It fits the inference-gateway/routing category because its core product sits between applications and model providers, deciding which model should handle each request and providing a shared control layer for inference.
Funding rounds
-
Seed
announced 2024-07-01
: 2300000 USD
; investors: defy.vc, Jeff Dean, Julien Chaumond, Zack Kass, Ion Stoica, Tom Preston-Werner, Scott Belsky, Jeff Weiner
; source: PR Newswire
; https://www.prnewswire.com/news-releases/not-diamond-launches-prompt-adaptation-an-agentic-system-for-multi-model-enterprise-ai-302459831.html
Company timeline
-
2026-09-07:
Not Diamond Code entered early access
(product)
. Not Diamond published documentation for Not Diamond Code, describing it as a model router built for long horizon coding agent workloads and early access use.
Source: Not Diamond Code docs.
https://code.notdiamond.ai/docs
-
2026-09-07:
HR snapshot
(hr)
. Not Diamond is a small San Francisco based team with engineering, research, revenue, and operations functions represented publicly. Public evidence shows an open invitation to apply, but not a clearly dated hiring surge.
Source: LinkedIn company page.
https://www.linkedin.com/company/notdiamond
-
2026-09-07:
Product snapshot
(product)
. Not Diamond is an intelligent AI model router and prompt optimization platform for developers building AI applications and coding agents.
Source: Not Diamond homepage.
https://www.notdiamond.ai
-
2026-09-01:
Interactive benchmarks methodology released
(product)
. Not Diamond published a methodology for evaluating model routing with interactive benchmarks and described tooling for clients to benchmark their own workloads.
Source: Not Diamond blog.
https://www.notdiamond.ai/blog/interactive-benchmarks-a-new-methodology-for-evaluating-model-routing
-
2026-07-01:
Current multi-function careers page active
(hr)
. Not Diamond’s live roles page currently lists open positions in AI devrel, backend/infra, enterprise technical account, forward deployed engineering, research, sales engineering, and security/DevOps.. Team size observed: 11-50 employees employees.
Source: Not Diamond careers page.
https://notdiamond.notion.site/
-
2026-01-20:
Prompt optimization reached general availability
(gtm)
. Not Diamond released prompt optimization GA, packaging the optimization workflow as a production-ready API capability and expanding the platform around model migration and prompt-level performance.
Source: Not Diamond Blog.
https://www.notdiamond.ai/blog/prompt-optimization-is-now-generally-available
Portkey
Website: portkey.ai
Portkey is an AI gateway and control-plane platform that gives developers and enterprises one unified layer for routing, observability, guardrails, and governance across multiple model providers.
Category: Inference gateways, routing & control.
Funding stage: Acquired.
Founded: 2023.
Country: United States.
Employees: 11-50 employees.
Landscape fit: It fits inference gateways/routing/control because Portkey sits between applications and multiple model providers, routes requests, and adds a shared production control layer for reliability, monitoring, and policy enforcement.
Funding rounds
-
Series A
announced 2026-02-19
: 15000000 USD
; investors: Elevation Capital, Lightspeed
; source: Portkey
; https://portkey.ai/blog/series-a-funding/
-
Seed
announced 2023-08-01
: 3000000 USD
; investors: Lightspeed India, Dev Khare, Manjot Pahwa, Sanjeev Sisodiya, Adit Parekh, Ankit Gupta, Shyamal Hitesh Anadkat, Manish Jindal, Jake Seid, Oliver Jay, Aakrit Vaish, Pranay Gupta, Sandeep Krishnamurthy, Gaurav Mandlecha, Miten Sampat
; source: Portkey
; https://portkey.ai/blog/building-a-full-stack-llmops-platform/
Company timeline
-
2026-09-07:
HR snapshot
(hr)
. Portkey appears to remain a small engineering organization in public signals, but current hiring is not clearly evidenced.
Source: LinkedIn company page.
https://www.linkedin.com/company/portkey-ai
-
2026-09-07:
Product snapshot
(product)
. Portkey provides an AI gateway and control plane for production AI systems. The current product surface includes routing, observability, guardrails, governance, MCP support, agent control, and model management, and the site now presents the product as the Prisma AIRS AI Gateway.
Source: Portkey homepage.
https://portkey.ai
-
2026-06-01:
Hiring across various engineering roles
(hr)
. Rohit Agarwal said Portkey was starting hiring across various engineering roles after Day 1 at Palo Alto Networks.. Team size observed: 11-50 employees employees.
Source: LinkedIn.
https://www.linkedin.com/posts/1rohitagarwal_portkey-was-a-crazy-0-to-1-journey-and-day-activity-7467873276300533760-yymg
-
2026-05-29:
Portkey acquired by Palo Alto Networks
(acquisition)
. Palo Alto Networks completed the acquisition of Portkey on 2026-05-29.. Acquirer: Palo Alto Networks. Deal amount: 140000000 USD. Status: completed
Source: Palo Alto Networks.
https://paloaltonetworks.gcs-web.com/news-releases/news-release-details/palo-alto-networks-completes-acquisition-portkey-secure-ai
-
2026-04-30:
Portkey acquired by Palo Alto Networks
(acquisition)
. Palo Alto Networks announced its intent to acquire Portkey to integrate it as the AI Gateway for Prisma AIRS. The company later announced the acquisition had closed on 2026-05-29.. Acquirer: Palo Alto Networks. Deal amount: 140000000 USD. Status: announced
Source: Palo Alto Networks.
https://paloaltonetworks.gcs-web.com/news-releases/news-release-details/palo-alto-networks-acquire-portkey-secure-rise-ai-agents
-
2026-04-29:
Hermes Agent integration shipped
(product)
. Portkey announced Hermes Agent support, exposing Portkey's OpenAI-compatible endpoint with 3,500+ models, logging, session-level cost tracking, budget limits, and MCP governance. ([new.portkey.ai](https://new.portkey.ai/announcements))
Source: Portkey announcements.
https://new.portkey.ai/announcements
Helicone
Website: helicone.ai
Helicone is an AI gateway and LLM observability platform for developers that provides a unified API plus logging, monitoring, routing, caching, and cost/performance analytics across model providers. ([ycombinator.com](https://www.ycombinator.com/companies/helicone?utm_source=openai))
Category: Inference gateways, routing & control.
Funding stage: Acquired.
Founded: 2023.
Country: United States.
Employees: 5.
Landscape fit: It fits the inference gateways, routing & control category because Helicone sits between applications and model providers, offering a shared production control layer for routing, retries, caching, monitoring, and model usage analytics. ([ycombinator.com](https://www.ycombinator.com/companies/helicone?utm_source=openai))
Funding rounds
-
Pre-Seed
announced 2023-04-01
: 1500000 USD
; investors: Coughdrop Capital, Y Combinator
; source: CB Insights
; https://www.cbinsights.com/company/helicone/financials
-
Seed
announced 2023-03-01
: 125000 USD
; investors: Y Combinator
; source: Y Combinator
; https://www.ycombinator.com/companies/helicone
Company timeline
-
2026-09-07:
HR snapshot
(hr)
. Helicone appears to be a team of 5 with no clear public hiring signal. The YC jobs page shows no open roles.
Source: Jobs at Helicone | Y Combinator.
https://www.ycombinator.com/companies/helicone/jobs
-
2026-09-07:
Product snapshot
(product)
. Helicone provides an AI gateway and LLM observability platform for developers. It lets teams route requests through a single API, log and inspect requests, monitor cost and latency, cache requests, manage prompts, apply fallbacks, and analyze LLM usage across providers.
Source: Documentation introduction.
https://docs.helicone.ai/introduction
-
2026-03-03:
Acquired by Mintlify and shifted to maintenance mode
(gtm)
. Helicone announced that it had been acquired by Mintlify and said the product would remain live in maintenance mode with ongoing fixes and model updates. ([helicone.ai](https://www.helicone.ai/blog/joining-mintlify?utm_source=openai))
Source: Helicone is joining Mintlify.
https://www.helicone.ai/blog/joining-mintlify
-
2026-03-03:
Helicone announced acquisition by Mintlify and team transition
(hr)
. Helicone said it had been acquired by Mintlify and that the Helicone team would be joining Mintlify in San Francisco.. Team size observed: 2-10 employees employees.
Source: Helicone blog.
https://www.helicone.ai/blog/joining-mintlify
-
2026-03-03:
Helicone acquired by Mintlify
(acquisition)
. Helicone announced that it had been acquired by Mintlify and that the Helicone team would join Mintlify in San Francisco. Helicone said its services would remain live in maintenance mode.
Source: Helicone Blog.
https://www.helicone.ai/blog/joining-mintlify
-
2025-11-26:
Claude Sonnet 4 and 4.5 1M context support added
(product)
. Helicone updated the AI Gateway so Claude Sonnet 4 and 4.5 used a 1M token context window by default across Anthropic API, Bedrock, and Vertex AI. ([helicone.ai](https://www.helicone.ai/changelog))
Source: Helicone Changelog.
https://www.helicone.ai/changelog
Requesty
Website: requesty.ai
Requesty is an AI gateway and LLM router that gives teams a single OpenAI-compatible endpoint to access 600+ models with routing, fallbacks, caching, governance, observability, and spend controls. ([uk.linkedin.com](https://uk.linkedin.com/company/requesty))
Category: Inference gateways, routing & control.
Funding stage: Seed.
Founded: 2023.
Country: United Kingdom.
Employees: 11-50.
Landscape fit: It fits this category because it sits between applications and model providers, centralizing model routing, fallback, policy, logging, and spend control for production inference. ([uk.linkedin.com](https://uk.linkedin.com/company/requesty))
Funding rounds
-
seed
announced 2025-09-26
: 3000000 USD
; investors: 20VC, Tapestry VC, Tiny Supercomputer Investment Company, Insiders
; source: Requesty official funding announcement
; https://www.requesty.ai/blog/requesty-raises-3m
Company timeline
-
2026-09-07:
HR snapshot
(hr)
. Requesty is hiring selectively, with public emphasis on founding engineering and growth roles in London. The company looks small and still in a build and scale phase.
Source: Requesty LinkedIn company page.
https://uk.linkedin.com/company/requesty
-
2026-08-20:
Product snapshot
(product)
. OpenAI-compatible AI gateway / LLM router with one endpoint for 600+ models, routing, fallbacks, caching, governance, observability, MCP Gateway support, spend controls, and EU data residency.
Source: Requesty homepage.
https://www.requesty.ai
-
2026-08-01:
Requesty positions startup program as self-serve distribution for affiliated startups
(gtm)
. Requesty introduced a startup program offering $1,000 in AI credits and six months free for startups affiliated with YC, 20VC, Tapestry, or TinyVC.
Source: Requesty startup program.
https://www.requesty.ai/startup
-
2026-08-01:
Requesty launches public pricing with free tier, PAYG, and enterprise plans
(gtm)
. The official pricing page introduced a free tier, 5% markup pay-as-you-go pricing, and enterprise plans with SSO, RBAC, custom SLAs, spend limits, budget caps, MCP Gateway, and EU data residency.
Source: Requesty pricing page.
https://www.requesty.ai/pricing
-
2026-07-01:
Requesty publishes named customer deployments and usage metrics
(gtm)
. The customer page lists deployments such as ZoomInfo, anwalt.de, Mozart AI, NotarBot, LOUPZ, ekkodale, and Soccerment, with usage and governance outcomes.
Source: Requesty customer page.
https://www.requesty.ai/customers
-
2026-07-01:
Requesty posts active founding-team roles in London
(hr)
. The careers page lists Founding Engineer, Founding DevOps/SRE Engineer, and Growth Lead roles and says the company is hiring its founding engineering team in London.
Source: Requesty careers page.
https://www.requesty.ai/careers
TrustedRouter
Website: trustedrouter.com
TrustedRouter is an AI routing company that gives developers one OpenAI compatible API for models from multiple providers, with controls for price, reliability, region, and privacy posture.
Category: Inference gateways, routing & control.
Funding stage: Seed.
Founded: 2026.
Country: United States.
Employees: unknown.
Landscape fit: It fits inference gateways, routing, and control because the product sits between applications and model providers, routes requests across models, and adds a shared trust and policy layer for production inference.
Funding rounds
-
seed
announced 2026-08-31
: 1250000 USD
; investors: Sam Lessin / Slow Ventures, Bill Tai, Linda Avey, George Xing, Peter Livingston / Unpopular Ventures, Katelyn Donnelly / Avalanche VC, Gert Lanckriet, Holmes Wilson, Tory Reiss, Daniel Imberman, Michael Staton, Jason Fang, Capitoria Ventures, Alexey Komissarouk, Henri Roussez, others
; source: TrustedRouter blog
; https://trustedrouter.com/blog/we-raised-1-25m-seed
Company timeline
-
2026-09-07:
HR snapshot
(hr)
. TrustedRouter appears to be a small founder led team with selective interest in systems, security, model serving, and evaluation talent. There is no clear public signal of broad hiring activity.
Source: Work on TrustedRouter.
https://trustedrouter.com/careers
-
2026-09-07:
Product snapshot
(product)
. TrustedRouter is an OpenAI compatible AI routing service with an attested prompt path, provider routing controls, and public trust evidence.
Source: About TrustedRouter.
https://trustedrouter.com/about
-
2026-09-04:
Status page goes live
(product)
. TrustedRouter published a public status page with live gateway health, latency, and incident history.
Source: TrustedRouter Status.
https://trustedrouter.com/status
-
2026-09-01:
Public launch on Product Hunt
(gtm)
. TrustedRouter launched publicly on Product Hunt with copy focused on attested routing, confidential routes, and provider failover.
Source: Product Hunt launch page.
https://www.producthunt.com/products/trustedrouter
-
2026-09-01:
HR snapshot
(hr)
. TrustedRouter is founder led and publicly based in Miami, but the supplied sources do not provide a reliable headcount. There is no clear public signal of broad hiring activity.
Source: About TrustedRouter.
https://trustedrouter.com/about
-
2026-09-01:
Product snapshot
(product)
. TrustedRouter offers an AI routing gateway and control layer with one OpenAI compatible API across many providers. The current surface includes provider routing, privacy floors, attestation evidence, signed receipts, catalogs of models and providers, batch jobs, video generation, and user provided models.
Source: About TrustedRouter.
https://trustedrouter.com/about
LiteLLM
Website: litellm.ai
LiteLLM is an open source AI gateway and proxy that gives teams one interface to route, control, and monitor requests across many model providers.
Category: Inference gateways, routing & control.
Funding stage: Seed.
Founded: 2023.
Country: United States.
Employees: 10.
Landscape fit: It fits inference gateways and routing control because the product sits between applications and model providers and manages model access, routing, and governance for production inference.
Funding rounds
-
seed
announced 2023-01-01
: 1600000 USD
; investors: Y Combinator, Gravity Fund, Pioneer Fund
; source: LiteLLM company page
; https://www.ycombinator.com/companies/litellm
Company timeline
-
2026-09-07:
HR snapshot
(hr)
. LiteLLM is a San Francisco based team with active hiring across technical and commercial roles. The live jobs page shows a small but broadening organization rather than a mature closed hiring posture.
Source: Jobs at LiteLLM.
https://www.ycombinator.com/companies/litellm/jobs
-
2026-09-07:
Product snapshot
(product)
. LiteLLM is an open source AI gateway and proxy with a Python SDK and FastAPI server for calling 100+ LLM providers in OpenAI format. It supports centralized authentication, spend tracking, budgets, routing, guardrails, logging, virtual keys, and an admin dashboard, with enterprise features for SSO, SCIM, OIDC or JWT auth, audit logs, secret management, and multi region deployment.
Source: LiteLLM documentation home.
https://docs.litellm.ai
-
2026-09-05:
Auto router compression splits by hop
(product)
. LiteLLM added separate compression settings for the routing classifier and the routed model call, cutting classifier costs further in testing.
Source: AutoRouter Per-Hop Compression: Cut LLM Classifier Costs Another 32% | liteLLM.
https://docs.litellm.ai/blog/auto-router-per-hop-compression
-
2026-09-01:
Jobs page shows broad hiring
(hr)
. LiteLLM's jobs page shows multiple current openings across sales, engineering, support, and developer relations.
Source: Jobs at LiteLLM.
https://www.ycombinator.com/companies/litellm/jobs
-
2026-08-27:
Town hall reports status dashboard
(product)
. LiteLLM said it shipped a public status dashboard and production Auto Router results in its August town hall recap.
Source: August Townhall Updates: Security, Stability, and Product | liteLLM.
https://docs.litellm.ai/blog/august-townhall-updates
-
2026-08-27:
Director of Security joins
(hr)
. LiteLLM said Oliver Jensen joined as Director of Security.
Source: August Townhall Updates: Security, Stability, and Product | liteLLM.
https://docs.litellm.ai/blog/august-townhall-updates
TrueFoundry
Website: truefoundry.com
TrueFoundry builds an enterprise AI gateway and deployment platform that gives teams a shared control layer for routing, governing, and observing model and agent traffic in production.
Category: Inference gateways, routing & control.
Funding stage: Series A.
Founded: 2021.
Country: India.
Employees: 51-200 employees.
Landscape fit: It fits this category because its core product is a control layer for inference and AI traffic, with routing, fallbacks, access control, monitoring, and governance for production deployments.
Funding rounds
-
series A
announced 2025-02-06
: 19000000 USD
; investors: Intel Capital, Peak XV Partners, Eniac Ventures, Jump Capital
; source: TechCrunch
; https://techcrunch.com/2025/02/06/intel-capital-fuels-truefoundrys-19m-funding-to-help-boost-ai-deployments-at-scale/
-
seed
announced 2022-09-19
: 2300000 USD
; investors: Sequoia India and Southeast Asia’s Surge, Eniac Ventures, Naval Ravikant, Dilip Khandelwal, Maneesh Sharma, Mike Boufford, Anthony Goldbloom
; source: official blog
; https://www.truefoundry.com/blog/announcing-our-seed-fund-message-from-the-founders
Company timeline
-
2026-09-07:
HR snapshot
(hr)
. TrueFoundry has an active hiring presence, but the supplied evidence only confirms a small number of current openings and does not support a firm team size estimate.
Source: Careers at TrueFoundry.
https://www.truefoundry.com/careers
-
2026-09-07:
Product snapshot
(product)
. TrueFoundry provides an enterprise AI gateway and deployment platform for production models, agents, and tool calls on customer owned cloud or Kubernetes infrastructure. The platform includes LLM routing, MCP gateway controls, agent governance, observability, access control, cost controls, and model deployment workflows.
Source: official website.
https://www.truefoundry.com
-
2026-08-26:
Staffbase case study published
(gtm)
. TrueFoundry highlighted Staffbase as a large production customer on its case studies page.
Source: TrueFoundry case studies page.
https://www.truefoundry.com/case-studies
-
2026-08-24:
NetApp case study highlighted
(gtm)
. TrueFoundry surfaced NetApp as a customer using its governed agent stack.
Source: Press Room.
https://www.truefoundry.com/press-room
-
2026-08-18:
TrueForge open sourced
(product)
. TrueFoundry open sourced TrueForge, a vendor neutral agent harness for production agents.
Source: TrueFoundry blog.
https://www.truefoundry.com/blog/engineering/trueforge-open-source-agent-harness
-
2026-08-18:
Solutions Architect role posted
(hr)
. TrueFoundry posted a Solutions Architect role focused on its largest AI Gateway customers and enterprise onboarding.
Source: LinkedIn job posting.
https://www.linkedin.com/jobs/view/solutions-architect-at-truefoundry-4446066924
Maxim AI
Website: getmaxim.ai
Maxim AI builds an enterprise AI evaluation, observability, and gateway platform that helps teams route, govern, test, and monitor production AI applications.
Category: Inference gateways, routing & control.
Funding stage: Seed.
Founded: 2023.
Country: United States.
Employees: 11-50.
Landscape fit: It fits the inference gateway and control layer category because its Bifrost product sits between applications and model providers, routes AI traffic, and adds governance, failover, and observability for production use.
Funding rounds
-
seed
announced 2024-06-18
: 3000000 USD
; investors: Elevation Capital, Undisclosed angel investors from Postman, Undisclosed angel investors from Chargebee, Undisclosed angel investors from Groww, Undisclosed angel investors from Razorpay, Undisclosed angel investors from Media.net
; source: Maxim AI official blog
; https://www.getmaxim.ai/blog/announcing-maxim-ais-general-availability-and-the-3m-funding-round-led-by-elevation-capital
Company timeline
-
2026-09-07:
HR snapshot
(hr)
. Maxim AI appears to be a small enterprise software team with limited public hiring. The company is led by its two founders and is recruiting selectively rather than running a broad hiring push.
Source: LinkedIn company page for Bifrost by Maxim AI.
https://www.linkedin.com/company/bifrost-ai-gateway
-
2026-09-07:
Product snapshot
(product)
. Maxim AI is an end to end platform for simulation, evaluation, observability, and governance of AI agents and applications. Its public product surface also includes Bifrost, an enterprise AI gateway for routing, guardrails, and traffic control.
Source: Platform Overview.
https://www.getmaxim.ai/docs/introduction/overview
-
2026-01-16:
Maxim updates product
(product)
. Maxim’s changelog highlights a logging and observability overhaul, an MCP gateway, and evals on file attachments.
Source: Maxim Updates - Maxim Blog.
https://www.getmaxim.ai/blog/tag/maxim-updates
-
2025-07-04:
Bifrost publicly released
(product)
. Maxim AI announced the public release of Bifrost, its open-source LLM gateway for high-throughput production AI systems, with routing, extensibility, and built-in observability.
Source: Maxim AI June 2025 updates.
https://www.getmaxim.ai/blog/maxim-ai-june-2025-updates/
-
2025-05-12:
Maxim publishes Clinc story
(gtm)
. Maxim said Clinc used its platform to benchmark and refine a conversational AI system for banking.
Source: Clinc customer story.
https://www.getmaxim.ai/blog/elevating-conversational-banking-clincs-path-to-ai-confidence-with-maxim
-
2025-05-01:
Maxim publishes Atomicwork story
(gtm)
. Maxim said Atomicwork used its platform to simulate user interactions before release and to support regression testing and validation of agent workflows.
Source: Atomicwork customer story.
https://www.getmaxim.ai/blog/scaling-enterprise-support-atomicworks-journey-to-seamless-ai-quality-with-maxim
Respan
Website: respan.ai
Respan is an AI engineering platform for LLM and agent products that provides routing, observability, evaluations, and control for production inference. ([respan.ai](https://www.respan.ai/docs/documentation/overview?utm_source=openai))
Category: Inference gateways, routing & control.
Funding stage: Seed.
Founded: 2023.
Country: United States.
Employees: 20-25 employees.
Landscape fit: It fits developer infrastructure because it sits between applications and model providers and gives developers a shared control layer for routing, tracing, evaluations, and fallback handling in production. ([respan.ai](https://www.respan.ai/docs/documentation/overview?utm_source=openai))
Funding rounds
-
seed
announced 2026-03-18
: 5000000 USD
; investors: Gradient Ventures, Y Combinator, Hat-Trick Capital, XIAOXIAO FUND, Antigravity Capital, Alpen Capital
; source: Respan official blog
; https://www.respan.ai/blog/respan-raises-5m-seed
Company timeline
-
2026-09-07:
HR snapshot
(hr)
. Respan is a 10-person team with focused hiring across engineering and technical go to market roles.
Source: Respan About.
https://www.respan.ai/about
-
2026-09-07:
Product snapshot
(product)
. Respan is a full stack AI engineering platform for LLM and agent products. It provides routing, observability, prompt management, evaluations, and red teaming.
Source: Respan docs.
https://www.respan.ai/docs/documentation/overview
-
2026-09-04:
Launch P1 privacy model
(product)
. Respan introduced P1 and said it now powers PII redaction across the platform.
Source: Respan LinkedIn.
https://www.linkedin.com/company/respan-ai
-
2026-09-01:
Customer scale on YC page
(gtm)
. YC says Respan is trusted by 100 plus AI startups and enterprise teams and processes 1B plus logs and 2T plus tokens each month.
Source: Y Combinator.
https://www.ycombinator.com/companies/respan
-
2026-09-01:
Software engineer AI opening
(hr)
. Respan posted a full time software engineer AI role focused on model routing, evaluations, prompt optimization, and trace analysis.
Source: Y Combinator jobs page.
https://www.ycombinator.com/companies/respan/jobs/9L4xvXt-software-engineer-ai
-
2026-06-11:
Launch AI gateway
(product)
. Respan launched an AI gateway with observability, evals, prompt optimization, and spend controls built in.
Source: Respan LinkedIn.
https://www.linkedin.com/posts/respan-ai_today-were-launching-respan-ai-gateway-activity-7470143275392241664-udTZ
Thesean
Website: thesean.ai
Thesean is an AI inference research lab incubated by Martian. Its first product, Ship, is a best-execution endpoint that aims to reduce frontier-model inference cost while preserving model capability and behavior.
Category: Inference gateways, routing & control.
Funding stage: .
Founded: 2026.
Country: United States.
Employees: unknown.
Landscape fit: It fits developer infrastructure because it sits between applications and model providers, routing and optimizing inference requests through an API rather than serving as a consumer AI app.
Company timeline
-
2026-09-07:
Claude Code integration becomes available
(product)
. Thesean documented how to configure Claude Code to use Ship through its Anthropic compatible Messages API.
Source: Claude Code integration.
https://docs.thesean.ai/integrations/claude-code
-
2026-09-07:
HR snapshot
(hr)
. Thesean shows a founder led, research heavy team profile, but the supplied evidence does not reveal headcount or a clear hiring push.
Source: Thesean AI homepage.
https://www.thesean.ai
-
2026-09-07:
Product snapshot
(product)
. Thesean provides Ship, a best execution endpoint for LLM requests that aims to reduce cost while keeping model behavior equivalent. The current public docs and homepage show OpenAI compatible Responses and Chat Completions, Anthropic compatible Messages, streaming, prompt caching, tool use, function calling, feedback driven optimization, and integrations for common coding tools.
Source: Thesean AI homepage.
https://www.thesean.ai
-
2026-09-01:
Cursor integration becomes available
(product)
. Thesean documented how to configure Ship inside Cursor using the Thesean API key and base URL.
Source: Cursor integration.
https://docs.thesean.ai/integrations/cursor
-
2026-09-01:
Claude Code integration becomes available
(gtm)
. Thesean documented a Claude Code setup that routes requests through Ship.
Source: Claude Code integration.
https://docs.thesean.ai/integrations/claude-code
-
2026-09-01:
Cursor integration becomes available
(gtm)
. Thesean documented a Cursor setup that routes requests through Ship.
Source: Cursor integration.
https://docs.thesean.ai/integrations/cursor
Concentrate AI
Website: concentrate.ai
Concentrate AI builds an LLM gateway that gives teams one API for multiple model providers, with routing, failover, logging, spend controls, and access governance for production inference. ([concentrate.ai](https://concentrate.ai/?utm_source=openai))
Category: Inference gateways, routing & control.
Funding stage: Pre Seed.
Founded: 2025.
Country: United States.
Employees: 2-10 employees.
Landscape fit: It fits the inference gateways and routing category because its core product sits between applications and model providers and controls request routing, fallback behavior, and spend governance. ([concentrate.ai](https://concentrate.ai/features/routing?utm_source=openai))
Funding rounds
-
pre-seed
announced 2026-06-10
: 5000000 USD
; investors: True Ventures, RRE Ventures
; source: CB Insights / RRE Ventures investment data
; https://www.cbinsights.com/investor/rre-ventures
Company timeline
-
2026-09-07:
HR snapshot
(hr)
. Concentrate AI looks like a small founder led startup with at least one GTM leader and a live engineering recruiting signal. The public evidence does not support a precise headcount.
Source: LinkedIn.
https://www.linkedin.com/jobs/view/ai-engineer-at-concentrate-ai-4368852571
-
2026-09-07:
Product snapshot
(product)
. Concentrate AI provides a unified LLM gateway and orchestration layer for production AI traffic. The platform includes one API for multiple model providers, automatic routing, failover, request logs, spend tracking, usage analytics, model access controls, RBAC, SSO or SAML, zero data retention, data redaction, team workspaces, key spend limits, audit logs, and pay as you go plus enterprise pricing.
Source: Concentrate.ai.
https://concentrate.ai
-
2026-07-01:
Supercode partnership
(gtm)
. A third party case study says Supercode partnered with Concentrate AI to power its AI requests through Concentrate's gateway.
Source: Supercode.
https://supercodeai.vercel.app/partnerships/concentrateai
-
2026-07-01:
GTM leader identified
(hr)
. A customer story from Meow identifies Zach Moskow as Head of GTM at Concentrate AI.
Source: Meow.
https://www.meow.com/customer-stories/why-concentrate-ai-runs-their-corporate-finances-on-meow
-
2026-06-10:
Launches out of stealth
(gtm)
. Concentrate AI launched out of stealth with its managed LLM gateway for production AI traffic. The launch said the company was already routing billions of tokens for customers and had raised more than $5M in pre-seed funding.
Source: Concentrate AI / Fortune press release.
https://fortune.com/press-releases/concentrate-ai-free-llm-gateway-ai-spend-2026-06-10/
-
2026-06-10:
Pre-seed
(funding)
. Amount: USD 5M. Investors: True Ventures, RRE Ventures
Source: CB Insights / RRE Ventures investment data.
https://www.cbinsights.com/investor/rre-ventures
On-device & local inference
Infrastructure built to run AI models directly on user-owned or edge hardware rather than relying primarily on remote cloud inference. These platforms make local models easier to deploy and operate while accounting for the constraints of the underlying device.
Ollama
Website: ollama.com
Ollama builds a local and cloud platform that makes it easy for developers and teams to run open AI models on their own hardware or nearby infrastructure.
Category: On-device & local inference.
Funding stage: Series B.
Founded: 2023.
Country: Canada.
Employees: 14.
Landscape fit: It fits on-device & local inference because its core product helps users download, run, and operate AI models on user-owned machines, with an emphasis on offline/local execution and hardware-aware deployment.
Funding rounds
-
Series B
announced 2026-07-09
: 65000000 USD
; investors: Theory Ventures
; source: TechCrunch: Ollama raises $65M Series B
; https://techcrunch.com/2026/07/09/popular-open-source-ai-developer-tool-ollama-raises-65m-grows-to-nearly-9m-users
Company timeline
-
2026-08-10:
Product snapshot
(product)
. Local and cloud platform for running open models with CLI/API access, desktop apps, model library access, official Python and JavaScript libraries, and many integrations.
Source: Ollama home page.
https://ollama.com
-
2026-07-09:
Team size observed
(hr)
. Public source shows a team size of 14 employees.
Source: TechCrunch: Ollama raises $65M Series B.
https://techcrunch.com/2026/07/09/popular-open-source-ai-developer-tool-ollama-raises-65m-grows-to-nearly-9m-users
-
2026-07-09:
Series B
(funding)
. Amount: USD 65M. Investors: Theory Ventures
Source: TechCrunch: Ollama raises $65M Series B.
https://techcrunch.com/2026/07/09/popular-open-source-ai-developer-tool-ollama-raises-65m-grows-to-nearly-9m-users
-
2026-06-11:
Ollama improves MLX performance on Apple Silicon
(product)
. Ollama's MLX engine was updated to deliver its highest Apple Silicon performance yet, including faster output, better quality, and support for NVIDIA NVFP4 imports.
Source: Ollama's highest performance on Apple Silicon yet with MLX · Ollama Blog.
https://ollama.com/blog/mlx-performance
-
2026-03-30:
Ollama previews MLX-powered Apple Silicon engine
(product)
. Ollama previewed an MLX-powered runtime on Apple Silicon aimed at faster execution and lower memory usage.
Source: Ollama is now powered by MLX on Apple Silicon in preview.
https://ollama.com/blog/mlx-performance
-
2026-01-23:
Ollama introduces ollama launch
(product)
. Ollama released a command to set up and run coding tools such as Claude Code, OpenCode, and Codex with local or cloud models without manual config.
Source: Ollama Blog: ollama launch.
https://ollama.com/blog/ollama-launch
LM Studio
Website: lmstudio.ai
LM Studio builds a desktop and developer platform for running open-source LLMs locally on user-owned hardware, with offline chat, model management, APIs, and tooling for app integration.
Category: On-device & local inference.
Founded: 2023.
Country: United States.
Employees: 2-10 employees.
Landscape fit: It fits the on-device/local inference category because its core product is explicitly designed to download, run, serve, and interact with LLMs locally rather than primarily through cloud inference.
Company timeline
-
2026-08-10:
Product snapshot
(product)
. Desktop and developer platform for running open-source LLMs locally on user-owned hardware, with offline chat, model management, local server APIs, SDKs, headless daemon deployment, and agentic tools for app integration.
Source: LM Studio Developer Docs.
https://lmstudio.ai/docs/developer
-
2026-08-02:
LM Studio adds DeepSeek V4 Flash cloud and local availability
(product)
. LM Studio Bionic added DeepSeek V4 Flash as both a local model download and a US-hosted cloud inference option with zero data retention by default.
Source: Run DeepSeek V4 Flash locally or in the cloud (US-hosted).
https://lmstudio.ai/blog/deepseek-v4-flash
-
2026-07-22:
LM Studio adds enterprise internal network model endpoint support
(product)
. LM Studio 0.4.20 added support for using this device’s models as an enterprise internal network model endpoint and expanded Bionic/LM Link behavior.
Source: LM Studio 0.4.20.
https://lmstudio.ai/changelog/lmstudio-v0.4.20
-
2026-07-16:
LM Studio introduces Bionic agent
(product)
. LM Studio launched Bionic, an AI agent for coding, research, and document workflows that can run locally or with cloud models.
Source: Introducing LM Studio Bionic: the AI agent for open models.
https://lmstudio.ai/blog/introducing-lm-studio-bionic
-
2026-07-01:
LM Studio lists GTM Enterprise Sales and Marketing Lead roles
(hr)
. The careers page includes GTM, Enterprise Sales and Marketing Lead openings in New York City.
Source: Careers at LM Studio.
https://lmstudio.ai/careers
-
2026-07-01:
LM Studio posts Application Software Engineer role
(hr)
. LM Studio posted a New York City opening for an Application Software Engineer to build polished UI for the desktop application and LM Link.
Source: Software Engineer, Application.
https://lmstudio.ai/careers/application-software-engineer
Conifer
Website: conifer.build
Conifer is an AI inference gateway that routes requests across local and cloud models, providing one interface, one account, and one bill for production inference.
Category: On-device & local inference.
Funding stage: Pre Seed.
Founded: 2025.
Country: United States.
Employees: 4.
Landscape fit: It fits this category because it sits between applications and model providers, managing routing, cost optimization, and control for inference workloads.
Company timeline
-
2026-07-10:
Privacy policy added explicit cloud, billing, and fleet governance controls
(product)
. On 2026-07-10, the privacy page spelled out cloud-lane routing, opt-in chat retention, account-based billing, and fleet governance with spend caps and local-only policy enforcement. ([conifer.build](https://www.conifer.build/privacy/?utm_source=openai))
Source: Conifer privacy policy.
https://www.conifer.build/privacy/
-
2026-07-10:
Privacy page expands the product into governed routing and team controls
(gtm)
. On 2026-07-10, the privacy page formalizes the cloud, live-data, account, billing, and routing-signal paths, and adds fleet governance features such as spend caps and a client-side local-only policy.
Source: Conifer Privacy page.
https://www.conifer.build/privacy/
-
2026-07-08:
YC S26 founder launch and public introduction
(hr)
. Michael Jeffords publicly introduced Conifer as a YC S26 company and said he and Charles Muehlberger were launching the product and welcoming inbound interest from teams.. Team size observed: 3 employees employees.
Source: LinkedIn / Y Combinator.
https://www.ycombinator.com/companies/conifer
-
2026-06-01:
Local-first router, scoring, and one-bill positioning became the product narrative
(product)
. By late June 2026, Conifer’s homepage framed the product as a local-first router that chooses the cheapest capable model, scores local models against cloud leaders, and charges one bill across local and cloud usage. ([conifer.build](https://www.conifer.build/?utm_source=openai))
Source: Conifer homepage and vision pages.
https://www.conifer.build/
-
2026-06-01:
Memory-aware context sizing fix shipped in Sage v1.2.4
(product)
. Conifer described a fix in Sage v1.2.4 that stopped the desktop path from forcing a full native context window and instead let the engine apply its memory-aware cap. ([conifer.build](https://www.conifer.build/research/fitting-big-models-to-small-memory/?utm_source=openai))
Source: Conifer research.
https://www.conifer.build/research/fitting-big-models-to-small-memory/
-
2026-06-01:
Research post uses benchmark-led content to prove local-first performance
(gtm)
. A June 2026 research article explains how Conifer fits large-model context windows to available memory, and it publishes measured throughput and memory-use comparisons to show why the local engine stays fast.
Source: Conifer Bench.
https://www.conifer.build/research/fitting-big-models-to-small-memory/
DiscreteStack
Website: discretestack.com
DiscreteStack builds private AI infrastructure for enterprises so they can run open models on their own servers with flat rate pricing and full control over data and governance. It says the company is built in Europe and is designed for on premises deployment. ([discretestack.com](https://discretestack.com/))
Category: On-device & local inference.
Funding stage: Seed.
Founded: 2025.
Country: Bulgaria.
Employees: 2-10.
Landscape fit: It fits on device and local inference because its core product is software that runs AI models on customer owned infrastructure instead of sending requests to a remote cloud API. The company also frames itself as an infrastructure layer with hardware native builds, scheduling, and governance for private deployments. ([discretestack.com](https://discretestack.com/))
Funding rounds
-
seed
announced 2026-08-17
: 800000 EUR
; investors: CleverPine Ventures, Milen Manev, Stoil Vasilev, several smaller investors
; source: DiscreteStack blog
; https://discretestack.com/blog/discretestack-raises-e800000-to-give-european-businesses-their-own-ai-infrastructure
Company timeline
-
2026-08-26:
DiscreteStack opens partner motion
(gtm)
. DiscreteStack published a partners page for application builders and AI integrators.
Source: DiscreteStack partners page.
https://discretestack.com/partners
-
2026-08-26:
Product snapshot
(product)
. Private AI infrastructure and runtime software for enterprises that deploy open models on their own servers, on premises, or in isolated environments.
Source: DiscreteStack home page.
https://discretestack.com
-
2026-08-23:
DiscreteStack documents setup workflow
(product)
. DiscreteStack published setup documentation showing live model names and a no prompt retention policy.
Source: DiscreteStack setup page.
https://setup.discretestack.com
-
2026-08-20:
DiscreteStack publishes pricing
(product)
. DiscreteStack published fixed pricing for its execution node product and framed it as flat rate infrastructure.
Source: DiscreteStack LLM product page.
https://discretestack.com/llm.html
-
2026-08-20:
DiscreteStack launches product page
(product)
. DiscreteStack published a product page that explains its private AI operating system and commercial terms.
Source: DiscreteStack LLM product page.
https://discretestack.com/llm.html
-
2026-08-17:
DiscreteStack says it serves first enterprise customers
(gtm)
. The funding announcement says the company already serves its first enterprise customers in Bulgaria and abroad.
Source: DiscreteStack blog.
https://discretestack.com/blog/discretestack-raises-e800000-to-give-european-businesses-their-own-ai-infrastructure