π¦ ZebraBench
Benchmark LLM inference endpoints right in your browser.
Available models
Columns
Load models to populate this table, then choose which ones to benchmark. Select any header to sort; blended pricing weights three input tokens for every output token.
What Speed Test 1 measures
Measures raw output-generation speed and charts one tokens-per-second value for each measured run.
Results will appear here after a Speed Test 1 run.
Token throughput per run
One chart per model with one tok/s bar per run. A dashed p50 line summarizes the completed runs.
Testing methodology
- TTFT runs from request dispatch to the first non-empty content or reasoning delta.
- Tok/s divides completion tokens by elapsed streaming time after that first delta.
- Thinking tokens count. When a model emits reasoning content, those tokens are included in both tok/s and TTFT. Because the server only reports a combined completion-token total, the reasoning share is estimated from streamed characters and flagged when present.
- End-to-end latency runs from request dispatch until the response stream closes.
- One warm-up request runs per model. It is excluded from throughput summaries but included in test time, token usage, and cost.
- Total run time is wall-clock time; parallel model times are not summed.
- Cost uses each model's displayed input and output price per million tokens; models without pricing remain unpriced.
- Each model gets a throughput chart with one tok/s bar per measured run; a dashed p50 reference line summarizes the completed runs.
- TTFT therefore represents short-prompt latency. Long-context prefill performance should be measured separately with controlled input sizes.
- Models run in parallel; measured runs for a given model remain serial. Use concurrency 1 for an isolated latency baseline.
- Model scheduling order is randomized and recorded in the export to reduce time-of-run bias.
-
Benchmark request
Live request template based on the current settings. The selected model is inserted when the test runs, and the API key is always redacted.
-
Sample request
Captured from a completed measured run. The API key is redacted.
Sample streaming response
Raw response headers and every decoded network chunk from the same measured run. Chunk labels are added by the benchmark.
Sample consolidated response
All generated deltas from the same measured run, consolidated in arrival order.
What Thinking Test 1 measures
Performance on a medium thinking task: given a bunch of records, pick the one with the highest score.
<id>|<name> line.
Results will appear here after a thinking benchmark run.
Testing methodology
- Prompt is generated per question from a deterministic seed. Each question presents a fresh randomized table of
{rows}rows with id, name, and score, then asks the model to return the id and name of the row with the highest score. - Scores are unique integers in a range larger than the row count, so there is always exactly one highest-scoring row with no ties to break.
- Names are deterministic labels in the form
user-1,user-2, and so on, so every row's name is unique within a question. - Row order is shuffled with a per-question seed, so the highest-scoring row is uniformly distributed across positions - the model cannot shortcut by reading the first or last row.
- Grading requires the model's entire final content to contain exactly one
<id>|<name>line (reasoning deltas are ignored and surrounding whitespace is tolerated). The id must use canonical positive-integer syntax and the name must match case-insensitively. Accuracy is correct answers / attempted measured questions; failed requests count as incorrect and non-compliant. - Format compliance is tracked separately - a run can be format-compliant (a parseable
<id>|<name>response) but incorrect, which separates "couldn't find the max" from "couldn't follow the format". - Reasoning tokens are estimated from the
reasoning_content/reasoningSSE deltas and shown alongside answer tokens (final content). The split is estimated - the server only provides the total. - TTFT and End-to-end latency are measured the same way as Speed Test 1 but on a per-question basis, so values reflect varying prompt sizes.
- Cost per correct answer = total cost Γ· correct runs. A model with high accuracy and low cost ranks first. Models with zero correct runs show "-".
- One warm-up question runs per model. It is excluded from accuracy, latency, and token summaries but included in test time, token usage, and cost.
- Total run time is wall-clock time; parallel model times are not summed.
-
Benchmark request
Live request template based on the current settings. A sample prompt with a real generated table is shown; the selected model is inserted when the test runs, and the API key is always redacted.
-
Sample request
Captured from a completed measured run. The API key is redacted.
Sample streaming response
Raw response headers and every decoded network chunk from the same measured run. Chunk labels are added by the benchmark.
Sample consolidated response
All generated deltas from the same measured run, split into reasoning and final content, plus the grading verdict.
What the Decode Test measures
Client-observed output generation speed from a short fixed prompt. Each model runs the configured number of tests at each configured output length (default: 100, 1000, and 4000 tokens); each length is summarized with p50/p90 percentiles.
Leaderboard Decode Speed (tok/s)
p50 per model by output length, sorted by the longest-length column. Select any column header to re-sort.
Decode Speed by output length
p50 tok/s per model across the configured output lengths. Hover a line or point for the full model id and values.
Run results
Results and per-run decode speeds will appear here after a Decode Test run.
Testing methodology
- Prompt is short and fixed so the test isolates output generation rather than prompt processing.
- Every configured output length per model β comma-separated in the form and defaulting to 100, 1000, and 4000 tokens β each repeated for the configured number of runs and reported as its own table row. The request sends
max_tokensequal to the length and, when enabled,min_tokensplusignore_eos: trueso the model cannot stop early. - Each row aggregates its runs with nearest-rank percentiles: decode speed p50; TTFT p50 and p90; total latency p50 and p90; plus p50 decode time, visible tokens, and reasoning tokens.
- Decode Speed = (visible output tokens β 1) Γ· seconds between the first and last visible output token. The first token is subtracted because it is not part of the decode interval.
- TTFT runs from request dispatch to the first visible (final content) token. Reasoning deltas are ignored for TTFT.
- Decode Time is the interval from the first to the last visible output token. Total Latency runs from request dispatch until the stream closes.
- Thinking is disabled when supported via
chat_template_kwargs: { enable_thinking: false }. If reasoning content still appears, it is excluded from decode speed and the run is marked reasoning required, with the reasoning token count shown separately. - Visible output tokens are the server completion-token total minus the reasoning share. When the endpoint reports a reasoning-token count (
completion_tokens_details.reasoning_tokensorreasoning_tokens), that exact value is used; otherwise the share is estimated from streamed reasoning vs. total output characters, and the run is flagged as estimated. - One warm-up request runs per model. It is excluded from decode-speed summaries but included in test time, token usage, and cost.
- Total run time is wall-clock time; parallel model times are not summed. Percentiles use nearest-rank selection across successful measured runs.
-
Benchmark request
Live request template based on the current settings. The selected model is inserted when the test runs, and the API key is always redacted.
-
Sample request
Captured from a completed measured run. The API key is redacted.
Sample streaming response
Raw response headers and every decoded network chunk from the same measured run. Chunk labels are added by the benchmark.
Sample consolidated response
All generated deltas from the same measured run, consolidated in arrival order.
What the Needle Test measures
Long-context needle retrieval with the document sized at a fill percent of each model's advertised context window (default 90%): each question hides one access-code entry at one of the configured percent positions (default 5%, 25%, 50%, 75%, 90%), and the accuracy matrix shows how retrieval holds across the document - the "lost in the middle" curve. Models run one at a time, and a confirmation dialog totals the planned requests, input tokens, and estimated cost before anything is sent.
Needle accuracy by position
One row per model, one column per needle position. Skipped cells (β) mean the model has no advertised context window to size the document against.
Run results
Results will appear here after a Needle Test run.
Testing methodology
- Prompt is generated per question from a deterministic seed: a fresh filler document of neutral log entries with exactly one access-code entry (the needle), so the model never sees the same document twice within a run.
- Documents fill a share of each model's context window (default 90%, from the Context fill % field), so every model is stressed at its own limit and models with larger windows get proportionally longer documents. Models without context-window metadata are skipped entirely - there is nothing to size against - and show as Skipped / β.
- Needle positions are the comma-separated percents (default 5, 25, 50, 75, 90). Placement is
round((lines β 1) Γ percent / 100), so 5% lands near the top and 90% near the bottom. The results table shows one row per model Γ needle position, and the Needle position column reports both the percent and the needle's approximate token depth into the document (~N tokens in). The matrix pivots accuracy into one row per model with one column per position. Task instructions stay fixed at the top and bottom of the prompt, so only the needle moves. - Status tracks each row live: Warming up, Running x/y while the combination's runs stream (the run in flight counts, so a row reads Running 1/3 as soon as its first run starts), Waiting while other positions run, then Completed x/y. Skipped x/y marks combinations that cannot run (a model without window metadata), Partial x/y / Failed x/y appear when requests fail, and Cancelled if you stop the run. The count is measured runs after the excluded warm-up. While a request's prompt is still uploading, that row's Status cell also shows a thin input progress bar (Input x%) - the request body is streamed in known-size chunks, so the bar tracks how much of the input has been sent; it disappears once the endpoint has accepted the request.
- Test time is a per-row column: the wall-clock time for that combination's measured runs (warm-up excluded). It ticks live from the row's first measured run, then freezes when the row completes; skipped and not-yet-started rows show "-".
- A styled confirmation dialog precedes every run: because documents fill most of each model's context window, one click can send millions of tokens. The dialog lists the planned requests and input tokens per model (warm-up included) with the estimated input cost from each model's pricing metadata, plus a highlighted totals block - and nothing is sent until you press Run test. Escape, Cancel, or clicking the backdrop backs out without a request. Output and reasoning tokens are not included in the estimate.
- Filler is neutral: no filler line mentions codes or sectors, and the needle is the only code-shaped string in the document, so retrieval cannot be shortcut by pattern-spotting.
- Grading requires the model's entire final content to be exactly one access-code line (reasoning deltas are ignored and surrounding whitespace is tolerated). The code must match the document entry case-insensitively. Failed requests count as incorrect and non-compliant, and count against their position.
- TTFT measures long-prompt prefill (request dispatch to the first generated token), shown as p50 (the typical run) and p90 (the tail - with few runs, nearest-rank p90 is simply the slowest run). Position does not change prompt length, so TTFT is comparable across positions.
- Effective input processing rate = p50 input tokens Γ· p50 TTFT per row. It includes network and queueing overhead - hence "effective" - and prefill must complete before any token, so the rate is a fair digest-speed figure. Higher is better.
- Thinking is left at each model's default behavior; the test sends no
chat_template_kwargs, so thinking models reason while retrieving and the reasoning-token columns show what that costs. - Reasoning tokens are estimated from the
reasoning_content/reasoningSSE deltas and shown alongside answer tokens. The split is estimated - the server only provides the total. - Cost per correct answer = total cost Γ· correct runs. A model with high accuracy and low cost ranks first. Models with zero correct runs show "-".
- One warm-up question runs per model. It is excluded from accuracy, latency, and token summaries but included in test time, token usage, and cost.
- Models run one at a time, never in parallel: a window-sized prefill is never contended by another model's traffic, so TTFT and the effective input rate stay comparable across models. Total run time is wall-clock time.
-
Benchmark request
Live request template based on the current settings, shown against an illustrative 128K window with the document truncated for readability; each model receives its own window-sized document when the test runs. The API key is always redacted.
-
Sample request
Captured from a completed measured run. Long documents are truncated head + tail and the API key is redacted.
Sample streaming response
Raw response headers and every decoded network chunk from the same measured run. Chunk labels are added by the benchmark.
Sample consolidated response
All generated deltas from the same measured run, split into reasoning and final content, plus the grading verdict.
What the Prefill Test measures
Long-prompt prefill speed across input sizes (default 10K, 50K, 100K, 250K, 500K, 1M tokens): each question hides one access-code entry near the end of the document (default 90%) and the effective input rate shows how fast the endpoint digests each size.
Run results
Results will appear here after a Prefill Test run.
Testing methodology
- Prompt is generated per question from a deterministic seed: a fresh filler document of neutral log entries with exactly one access-code entry (the needle), so the model never sees the same document twice within a run.
- A styled confirmation dialog precedes every run: mega-token prompts mean one click can send millions of tokens. The dialog lists the planned requests and input tokens per model (warm-up included) with the estimated input cost from each model's pricing metadata, plus a highlighted totals block - and nothing is sent until you press Run test. Escape, Cancel, or clicking the backdrop backs out without a request. Output and reasoning tokens are not included in the estimate.
- Input sizes are the comma-separated token counts (default 10K, 50K, 100K, 250K, 500K, 1M), tested smallest first. Sizes are directly comparable across models. Combinations whose size exceeds a model's advertised context window are skipped without sending a request and show as Skipped / β; models without window metadata are always sent.
- The needle position is fixed (default 90%) so the only variable is the input size. Placement is
round((lines β 1) Γ percent / 100), near the end of the document. The results table shows one row per model Γ input size, with the Needle position column reporting the percent and the needle's approximate token depth (~N tokens in). Task instructions stay fixed at the top and bottom of the prompt, so only the size moves. - Status tracks each row live: Warming up, Running x/y while the size's runs stream (the run in flight counts, so a row reads Running 1/3 as soon as its first run starts), Waiting while other sizes run, then Completed x/y. Skipped x/y marks sizes that exceed a model's advertised context window, Partial x/y / Failed x/y appear when requests fail, and Cancelled if you stop the run. The count is measured runs after the excluded warm-up. While a request's prompt is still uploading, that row's Status cell also shows a thin input progress bar (Input x%) - the request body is streamed in known-size chunks, so the bar tracks how much of the input has been sent; it disappears once the endpoint has accepted the request.
- Test time is a per-row column: the wall-clock time for that size's measured runs (warm-up excluded). It ticks live from the row's first measured run, then freezes when the row completes; skipped and not-yet-started rows show "-".
- Filler is neutral: no filler line mentions codes or sectors, and the needle is the only code-shaped string in the document, so retrieval cannot be shortcut by pattern-spotting.
- Grading requires the model's entire final content to be exactly one access-code line (reasoning deltas are ignored and surrounding whitespace is tolerated). The code must match the document entry case-insensitively. Failed requests count as incorrect and non-compliant, and count against their size.
- TTFT measures long-prompt prefill (request dispatch to the first generated token) at each input size, shown as p50 (the typical run) and p90 (the tail - with few runs, nearest-rank p90 is simply the slowest run).
- Effective input processing rate = p50 input tokens Γ· p50 TTFT per size row - the headline metric. It includes network and queueing overhead - hence "effective" - and prefill must complete before any token, so the rate is a fair digest-speed figure. Higher is better, and the curve across sizes shows whether the endpoint keeps up as prompts grow.
- Thinking is left at each model's default behavior; the test sends no
chat_template_kwargs, so thinking models reason while retrieving and the reasoning-token columns show what that costs. - Reasoning tokens are estimated from the
reasoning_content/reasoningSSE deltas and shown alongside answer tokens. The split is estimated - the server only provides the total. - Cost per correct answer = total cost Γ· correct runs. A model with high accuracy and low cost ranks first. Models with zero correct runs show "-".
- One warm-up question runs per model (the smallest size). It is excluded from accuracy, latency, and token summaries but included in test time, token usage, and cost.
- Total run time is wall-clock time; parallel model times are not summed.
-
Benchmark request
Live request template based on the current settings, shown for the smallest input size with the document truncated for readability; the selected model receives every size when the test runs. The API key is always redacted.
-
Sample request
Captured from a completed measured run. Long documents are truncated head + tail and the API key is redacted.
Sample streaming response
Raw response headers and every decoded network chunk from the same measured run. Chunk labels are added by the benchmark.
Sample consolidated response
All generated deltas from the same measured run, split into reasoning and final content, plus the grading verdict.
What the Cache Test measures
Prompt-cache effectiveness with a simple trivia question asked three ways: a short prompt (the bare question, below the cacheable minimum), a padded prompt (the same question with filler that reaches the configured size), and a nonce-busted control (the padded prompt with a fresh nonce prepended to every request, so the prefix cache can never match). Each variant sends one cold request then the configured number of repeats. The server-reported cached-token split shows what the endpoint actually reused, next to the TTFT drop and the billed-cost drop - and the busted rows show what no caching at all looks like.
TTFT: cold vs warm
One chart per model: the cold request's TTFT (muted) against the warm p50 (accent) for each variant, on a log scale. The gap is the prefill the cache skipped - padded should drop hard, busted should not move (by design), and short sits below the cacheable minimum.
Run results
Results will appear here after a Cache Test run.
Testing methodology
- One fixed question, three variants. Every request asks "What is the capital of France?" The short variant is the bare question - a handful of tokens, deliberately below the minimum prefix most providers cache. The padded variant pads the same question with a seeded filler log sized to the configured payload (default 50K tokens), well above that minimum. The busted control sends the same padded prompt with a fresh nonce as the very first line of every request - and since prefix caches match from the first token, a differing first line guarantees a miss no matter what follows.
- The request sequence per variant is 1 cold + N repeats. The cold request is the cache prime and the prefill baseline; the repeats (default 5 per variant) are sent back to back immediately after it. There is no excluded warm-up: every cold request is itself a measured data point. Short and padded resend byte-identical bodies; busted rebuilds its body around a new nonce every time.
- The busted rows are the no-cache control. Their cold/warm split is meaningless by design - every request is a guaranteed miss - so expect ~0% cache hit, ~0% TTFT reduction, and ~0% cost savings there. The contrast between the padded and busted rows of the same model is the cleanest read of the cache effect: identical prompt content, one cacheable and one not.
- The padding is regenerated from a fresh seed per benchmark run, so re-running the test does not measure a cache warmed by the previous run. It is padding, not a cache-busting nonce - nothing random is injected into individual requests of the short or padded variants, and within a run every one of their requests sends the same bytes. A cold request can still hit an unrelated cache (for example shared provider infrastructure), so treat cold numbers as "first request", not a guaranteed miss.
- Request bytes are identical within a variant: the same body object is serialized for the cold request and every repeat, so nothing but the prefix cache can explain any difference between them.
- Cache hit % is server-reported, not inferred: warm runs divide the usage
prompt_tokens_details.cached_tokens(or the endpoint'scached_tokensshorthand) byprompt_tokens. Endpoints that do not report the split show "-". - TTFT reduction = (cold TTFT β warm TTFT p50) Γ· cold TTFT. TTFT runs from request dispatch to the first generated token, so it contains the prefill the cache is supposed to skip; with few runs, nearest-rank p90 is simply the slowest warm run.
- Cost per request = uncached prompt tokens Γ input price + cached tokens Γ cached-input price, from the model's catalog pricing. Cost savings compares the warm cost with the cold cost; models without cached pricing metadata bill the cached share at the full input price (fallback), so their savings read 0% until a cached price is added to the catalog.
- Warm accuracy grades each repeat's answer (it must mention Paris), confirming a cached prefill does not corrupt the reply.
- Prefix caching caveats apply: caches are strictly prefix-based, live for minutes (not forever), and many endpoints cache only prefixes above a minimum length. This test measures the happy path - one stable payload, immediate reuse - not partial-prefix overlap or TTL decay.
- Thinking is left at each model's default behavior; the test sends no
chat_template_kwargs, matching the long-context tests, so reasoning models may spend reasoning tokens before the first visible token. TTFT counts the first generated token of any kind. - Models run one at a time so each cold prefill is uncontended and TTFT stays comparable. Total run time is wall-clock time.
-
Benchmark request
Live request template based on the current settings; the selected model is inserted when the test runs, and the API key is always redacted. The identical body is sent for the cold request and every warm repeat.
-
Sample request
Captured from a completed warm run. Long documents are truncated head + tail and the API key is redacted.
Sample streaming response
Raw response headers and every decoded network chunk from the same warm run, including the usage chunk that reports the cached-token split. Chunk labels are added by the benchmark.
Sample consolidated response
All generated deltas from the same warm run, consolidated in arrival order, plus the grading verdict.