Performance statistics

How fast is just one
Cortega gateway?

Cortega is built to compete at the top of AI gateway performance benchmarks. Based on the public benchmark patterns used in this market and the numbers we have measured so far, we believe Cortega belongs in the fastest tier of AI gateways. We tested a single Cortega gateway instance to show the cost of putting governance infrastructure in the request path.

The numbers below are for one gateway instance. Cortega is designed to run as a distributed layer of gateways that is managed as one single unified entity.

Our view is simple: if a buyer wants to compare AI gateways, bring the benchmark. Throughput, latency, cost per million requests, horizontal scale, reporting load, guardrails, policy enforcement, and production readiness all matter. Cortega was built for that comparison.

52,000requests per second on one 8 vCPU gateway, at 100% success
~2 msestimated median gateway overhead at that rate
100%success at every rate the tested boxes could carry
~$0.006per million requests for the three-node c9g setup
How we tested

Separate load generator. One gateway. One upstream provider

The gateway and the control plane ran on separate nodes.

The upstream provider was a mock server with a fixed 60 ms delay. That keeps provider latency constant, so the thing left to measure is what the gateway adds. Each test ran for 30 seconds at a fixed request rate.

Results

What the gateway added.

The overhead below is the median end-to-end latency minus the 60 ms mock upstream. The table keeps the raw end-to-end latency beside it for context.

vCPU $/hr Requests/sec End-to-end p50 Gateway overhead Success
2$0.0875,00061.7 ms~2 ms100%
2$0.0878,00061.9 ms~2 ms100%
2$0.08712,00064.8 ms~5 ms100%
4$0.17415,00060.6 ms~1 ms100%
4$0.17425,00061.9 ms~2 ms100%
8$0.34840,00062.0 ms~2 ms100%
8$0.34848,00062.7 ms~3 ms100%
8$0.34852,00064.9 ms~5 ms100%
8$0.34853,00069.3 msout of headroom99.97%
Reading the numbers

A small instance carried high request volume with low median overhead.

At loads each node could carry, the gateway added a few milliseconds. A 2 vCPU node held 12,000 requests per second and added about 5 ms. A 4 vCPU node held 25,000 requests per second and added about 2 ms. An 8 vCPU c9g.2xlarge held more than 52,000 requests per second, still at 100% success.

Throughput scaled with the cores. Doubling from 4 to 8 vCPUs roughly doubled the rate, from about 25,000 to more than 52,000 requests per second.

Failure case: The 8 vCPU box found its ceiling at 53,000 requests per second, where success slipped to 99.97%.

Cost

The full test setup was inexpensive to run.

The two 2 vCPU nodes together cost about 26 cents an hour. At 12,000 requests per second, that is about 43 million requests an hour, or roughly $0.004 per million requests for the full setup.

The two 4 vCPU setup lands at about the same cost per million at 25,000 requests per second. Counting only the gateway node, it is closer to $0.002 per million requests.

Scope

What these numbers represent.

These tests isolate the gateway path by holding upstream latency constant at 60 ms. That makes the added gateway overhead easier to read.

The overhead figures are estimates by subtraction at the median. We keep the raw latency numbers visible so the test can be read without hiding the setup.

The main takeaway is simple: one small Cortega gateway moved tens of thousands of requests per second with a few milliseconds of median overhead in this setup.

If you are comparing Cortega with other AI gateways, read our open-source AI gateway evaluation. It covers LiteLLM-style gateways, architecture, database and reporting implications, and how Cortega's gateway design differs.

Scale-out direction

One gateway is the starting point.

While this shows vertical scaling is effective, Cortega is also able to scale horizontally. The cost per million requests will remain the same but customers can benefit from higher redundancy and colocating gateways close to traffic originationss.

That matters for customers because AI governance should be able to start small, then expand with traffic, teams, agents, and providers.

Want to talk through the numbers?

We can walk through the setup, the request path, and what these numbers mean for your environment.

Request a conversation