AI platforms compared: what each one is actually built for
Search for AI platforms and you get a list of company names. That list is close to useless, because the word covers at least five different products, each solving a different bottleneck. An inference engine, a GPU cloud, a Kubernetes serving layer, an agent runtime and a machine learning lifecycle platform all call themselves platforms. Mostly they don’t compete. They stack.
So the first decision isn’t which vendor to pick. It’s which layer your problem actually sits on. If your GPUs are half idle, an agent runtime won’t fix that. If your agents are leaking state between users, a faster inference engine won’t fix that either. Picking the wrong layer costs more than picking the wrong company inside one.
What follows reads each platform’s own documentation and takes the design constraint it names for itself. Every quote below comes from a page we opened in September 2026. Where a figure is a vendor’s own claim rather than an independent measurement, we say so. If you want the billing question instead, our comparison of generative AI platforms by API, marketplace or open weights covers that, and this piece covers what sits underneath it.
One word, five jobs
The layers differ by the resource they’re trying to protect. An inference engine protects GPU memory. A GPU cloud protects the physical machine. A Kubernetes serving layer protects your budget during the quiet hours. An agent runtime protects the boundary between one user’s session and another’s. A lifecycle platform protects the record of what you shipped.
| Layer | Named examples | What it’s built to protect | What it won’t do |
|---|---|---|---|
| Inference engine | vLLM, SGLang | KV cache memory on the card | Get you a card, or bill anyone |
| GPU cloud | CoreWeave Kubernetes Service | The machine and its network path | Decide how you serve a model |
| Kubernetes serving layer | KServe, Ray Serve | Spend during idle hours | Make a single request faster |
| Agent runtime | Bedrock AgentCore, Gemini Enterprise Agent Platform | Session isolation and tool permissions | Improve the model’s reasoning |
| Lifecycle platform | SageMaker AI, MLflow | The record of models and experiments | Serve frontier models fastest |
Most vendors sell across two or three of these rows at once, which is why directory articles blur them. AWS alone sells a GPU cloud, a lifecycle platform and an agent runtime. That doesn’t make them one product. It makes them three purchases on one invoice.
Inference engines exist because GPU memory was being wasted
The bottom of the serving stack is the least glamorous and the best documented. vLLM describes itself as “Easy, fast, and cheap LLM serving for everyone”, and lists “State-of-the-art serving throughput” first among its features. The mechanism is named right after it: “Efficient management of attention key and value memory with PagedAttention”, plus “Continuous batching of incoming requests, chunked prefill, prefix caching”.
That’s a memory management product, not a model product. The paper behind it explains why. For the 13B parameter OPT model, the authors calculate that the key-value cache for a single token needs 800 KB, so one request generating up to 2,048 tokens can hold 1.6 GB of cache. Older serving systems reserved a contiguous block at the maximum length, which meant most of that space sat unused.
Indeed, our profiling results in Fig. 2 show that only 20.4% – 38.2% of the KV cache memory is used to store the actual token states in the existing systems.
Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention
Which means roughly three fifths to four fifths of the most expensive memory in the building was doing nothing. The paper reports that fixing it improved throughput “by 2-4× with the same level of latency” against FasterTransformer and Orca, the state of the art at the time. On OPT-13B specifically, the authors report vLLM handling 2.2 times more concurrent requests than Orca with oracle knowledge of output lengths, and 4.3 times more than the version that assumes maximum length. Those are the authors’ own measurements, published in 2023, and the baselines have moved since.
SGLang attacks the same layer from a different angle. Its docs say it’s “Designed for low-latency, high-throughput inference with RadixAttention, prefix caching, and multi-GPU parallelism”, and it claims real deployment scale: it “powers large-scale production deployments, generating trillions of tokens each day across more than 400,000 GPUs worldwide”, hosted under the non-profit LMSYS. We couldn’t verify that GPU count independently, and no vendor publishes the methodology behind it.
The practical point is that both projects are free software, and neither gives you a machine to run it on. That’s the next layer up, and it’s where the money goes.
GPU clouds sell the machine, and they’re specific about the machine
A GPU cloud competes on things a general cloud treats as implementation detail. CoreWeave’s Kubernetes Service documentation leads with the header claim of “High-performance managed Kubernetes on bare metal with DPU isolation and per-cluster VPCs”. The sentence that follows is the whole pitch: “CKS runs Kubernetes directly on bare metal Nodes, without a hypervisor. Customer clusters don’t run Virtual Machines.”
That’s a deliberate subtraction. A hypervisor is the layer that lets one physical server pretend to be several, and removing it costs you flexibility to buy back predictability. CoreWeave then adds NVIDIA BlueField Data Processing Units to each node “to offload processing tasks”, so network and storage work moves off the host CPU. The stated goal in the docs is “granular control, high performance, enhanced security, and high reliability, as well as high visibility into cluster metrics”.
The catch is that opinionated infrastructure is opinionated about your tooling too. The same page carries a blunt warning: “Do not install the NVIDIA GPU Operator on CKS clusters.” Doing it conflicts with the platform-managed deployment and isn’t supported. That’s a reasonable trade, and it’s also the kind of constraint that never appears in a feature comparison table. We worked through the cost side of renting hardware in our piece on whether to buy a seat or run a GPU.
The Kubernetes serving layer is built for the hours nobody uses it
Between the engine and the cloud sits a layer that does something neither of them does: it turns capacity off. KServe describes itself as a “Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes”, and calls itself “the open-source standard for self-hosted AI”. A banner across its homepage says it has joined the CNCF, and its footer lists it as “a Series of LF Projects, LLC”.
Read its feature list and the theme is operations, not speed. It offers “Scale-to-zero on expensive resources when not in use”, “Canary rollouts, pipelines, and ensembles with InferenceGraph”, and a “Standardized inference protocol for model servers with request/response APIs, supporting both predictive and generative models”. None of that makes a single token arrive sooner. All of it decides what your bill looks like on a Sunday.
The stacking is explicit on the page. Under generative AI, KServe lists its “Optimized Backends” as “vLLM and llm-d for high-performance LLM serving”. So the Kubernetes layer runs the engine from two sections ago rather than replacing it, which is the clearest evidence that these categories aren’t rivals.
Ray Serve sits in the same row and is unusually candid about the limits. Its docs call it “a scalable model serving library for building online inference APIs”, and highlight “fractional GPUs so you can share resources and serve many machine learning models at low cost”. Then, comparing itself to AWS SageMaker, Azure ML and Google Vertex AI, it says the quiet part out loud.
Ray Serve is not a full-fledged ML Platform. […] Ray Serve primarily focuses on model serving and providing the primitives for you to build your own ML platform on top.
Ray Serve documentation, comparing itself to SageMaker, Azure ML and Vertex AI
The same page also states that Ray Serve “doesn’t perform any model-specific optimizations to make your ML model run faster”. So it loses a raw throughput contest against a dedicated engine on purpose, and it says so in its own documentation. That honesty is more useful than a benchmark, because it tells you the axis on which the project isn’t trying to win.
Agent platforms are mostly not about the model
The newest layer is the one whose name misleads people the most. An agent platform sounds like it makes agents smarter. Read the component list and it plainly doesn’t. Amazon Bedrock AgentCore describes itself as “an agentic platform for building, deploying, and operating highly effective agents securely at scale using any framework and foundation model”, and it ships thirteen modular services. Two of them host the reasoning loop. The other eleven handle identity, isolation, memory, tooling, policy, evaluation and payments.
| Concern | AWS Bedrock AgentCore | Gemini Enterprise Agent Platform |
|---|---|---|
| Hosting the loop | Runtime, Harness | Agent Runtime |
| Remembering across turns | Memory | Sessions, Memory Bank |
| Running untrusted code | Code Interpreter, Browser | Code Execution, Computer Use |
| Authenticating the agent | Identity, Gateway | Not named on the page we read |
| Constraining tool calls | Policy, Registry | Not named on the page we read |
| Judging output quality | Evaluations, Optimization, Observability | Example Store, Evaluation Service, Feedback service |
| Paying for external calls | Payments | Not named on the page we read |
Isolation is where the engineering actually went. AgentCore Runtime states that “each user session runs in a dedicated microVM with isolated CPU, memory, and filesystem resources”, and that after a session ends “the entire microVM is terminated and memory is sanitized”. Sessions “run for up to 8 hours on microVMs, or up to 14 days on Instances”, and the runtime “can process 100MB payloads”. Those are architecture decisions about blast radius, not about intelligence.
Google frames the same problem in the same order. Its overview opens with “Bringing AI agents into production requires a high-performance runtime and a systematic approach to continuous improvement”, then lists serverless efficiency, context management through Sessions and a Memory Bank, continuous quality improvement through an Example Store and Evaluation Service, and “Secure sandbox execution”. Both vendors support outside frameworks, including LangGraph and LlamaIndex, so neither is selling you an agent. They’re selling you the cage it runs in.
That matters because the failures these platforms address are the ones that stop deployments, not the ones that lose benchmarks. We looked at where those projects stall in our piece on what enterprise AI actually gets into production.
Lifecycle platforms keep the receipts
The oldest category has been quietly renamed around the newest one. Amazon SageMaker AI calls itself “a fully managed machine learning (ML) service” where teams “build, train, and deploy ML models into a production-ready hosted environment”. AWS renamed the original SageMaker to SageMaker AI on 3 December 2024, and now uses the bare SageMaker name for “a unified platform for data, analytics, and AI” that bundles the lakehouse, governance, SQL analytics and Bedrock alongside it.
MLflow shows the same category stretching to fit generative models. Its documentation splits in two. The machine learning half offers “experiment tracking, model packaging, registry management, and deployment”. The other half offers “tools for LLM and agent observability, prompt management, foundation model deployment, and evaluation frameworks”. Same product, two audiences, one underlying idea: write down what you ran so you can answer for it later.
That’s the value nobody buys until an auditor asks. It’s also the layer that ages best, because a registry entry outlives whichever model generation filled it. If the wider stack below this is unfamiliar, our walkthrough of the six layers from chips to agents sets out the hardware underneath.
Which layer loses, and what would change the answer
On this evidence the Kubernetes serving layer loses the contest it’s most often entered into. Teams reach for KServe or Ray Serve expecting faster inference, and neither claims that. Ray Serve says it does no model-specific optimisation at all. What that layer buys is scale-to-zero, canary rollouts and a stable protocol, so it pays off when traffic is spiky and costs you complexity when it isn’t.
The lifecycle platforms lose a narrower contest. They’re built for a pipeline that trains and registers a model you own, and a team calling a frontier API doesn’t have that pipeline. The tracking and evaluation halves still apply, which is exactly the half MLflow has been extending.
Here’s what we couldn’t verify. We found no independent head-to-head benchmark run on identical hardware across vLLM, SGLang and the managed runtimes, so every throughput figure above is a project’s own. SGLang’s 400,000 GPU claim has no published methodology behind it. And the PagedAttention numbers come from 2023 against baselines that have been rewritten since, so treat the 2-4× as the reason the technique spread rather than as today’s margin.
Two things would change this reading. If a neutral party published a serving benchmark on fixed hardware with fixed request traces, the engine comparison would stop running on vendor assertion. And if the agent runtimes start shipping model-side improvements rather than isolation and policy, the layer would stop being infrastructure and start being a product you’d choose a model for. Until then, the useful question when a vendor says platform is simply which of these five rows they’re standing on.
Sources: vLLM documentation; Efficient Memory Management for Large Language Model Serving with PagedAttention; SGLang documentation; CoreWeave Kubernetes Service documentation; KServe; Ray Serve documentation; Amazon Bedrock AgentCore developer guide and AgentCore Runtime; Amazon SageMaker AI documentation; Gemini Enterprise Agent Platform overview; MLflow documentation. All read September 2026.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
