Model comparison

Compare the model. Then compare the system.

Kendr Intelligent is not another foundation model. It is a routing mode that chooses an eligible model for each task. This page compares that system with leading fixed models across cost, context size, output limits, efficiency mechanics, and workload fit.

Published pricing in USDPer 1 million tokensUpdated July 18, 2026Official provider sources
At a glance

Four different choices

The fixed models below are leading high-capability options from their providers. Kendr Intelligent sits above model selection and can use different eligible routes as task requirements change.

Adaptive system

Kendr Intelligent

kendr-intelligent
Price
Route dependent
Context
Up to 1M
Selection
Automatic
OpenAI flagship

GPT-5.6 Sol

gpt-5.6-sol
Price
$5 / $30
Context
1.05M
Selection
Fixed model
Anthropic Opus

Claude Opus 4.8

claude-opus-4-8
Price
$5 / $25
Context
1M
Selection
Fixed model
Google preview

Gemini 3.1 Pro

gemini-3.1-pro-preview
Price
$2 / $12*
Context
1.05M
Selection
Fixed model
Specifications

Cost, context, and efficiency

Published specifications are useful, but they describe capacity rather than guaranteed answer quality. Efficiency also depends on how much context is sent, which eligible route is selected, and whether a request retries or falls back.

CriteriaKendr IntelligentGPT-5.6 SolClaude Opus 4.8Gemini 3.1 Pro
Product typeMulti-model routerSelects an eligible route for the task.Foundation modelOne OpenAI model per request.Foundation modelOne Anthropic model per request.Foundation modelOne Google model per request.
Input / 1M tokensAdaptive + 5%Depends on the selected route and optimized context; Kendr adds one fixed 5% model-cost markup.$5.002× input price above 272K input tokens.$5.00Standard API rate across the full 1M context.$2.00$4.00 when the prompt exceeds 200K tokens.
Output / 1M tokensAdaptiveActual routed usage is charged to the Kendr plan balance.$30.001.5× output price above 272K input tokens.$25.00Thinking tokens are billed as output.$12.00$18.00 when the prompt exceeds 200K tokens; includes thinking.
Context windowUp to 1MRouter filters out routes that cannot satisfy required context.1,050,000Published context window.1,000,000Default context window.1,048,576Maximum input token limit.
Maximum outputRoute dependentUses the selected model's output limit.128,000Maximum output tokens.128,000Maximum output tokens.65,536Maximum output tokens.
Cost efficiencyTask-level optimizationBalances predicted quality, normalized cost, and latency among eligible routes, with bounded fallback when needed.Application-managedUse caching, batch, a smaller GPT tier, or shorter prompts to reduce cost.Application-managedUse prompt caching, batch, or a Sonnet/Haiku tier to reduce cost.Tiered by prompt sizeBatch, caching, and Flex can reduce standard inference cost.
Token efficiencyBuilt-in optimizerReduces irrelevant context before routing while preserving task instructions.Developer controlledThe application decides what context to send.Developer controlledContext management and caching are configured by the application.Developer controlledCaching and context composition are configured by the application.
Routing auditBuilt-in receiptReturns the selected public alias, route reason, versions, fallback outcome, latency, usage, and settled cost.Application-managedCapture provider response and usage metadata in your application.Application-managedCapture provider response and usage metadata in your application.Application-managedCapture provider response and usage metadata in your application.
Best fitMixed workloadsTeams that want automatic cost/capability routing across research, code, automation, and everyday tasks.Complex professional workHigh-end reasoning, coding, and broad tool use on one model.Agentic enterprise workComplex coding, long-running tasks, and adaptive reasoning.Multimodal agentsSoftware engineering, grounded workflows, audio, video, PDFs, and Google tools.

Swipe horizontally to view every model.

Cost example

Estimate one request

Change the token counts to compare direct provider inference cost. Tool calls, search, storage, caching, batch discounts, taxes, and plan pricing are not included.

Request size

Use token counts from a representative production request.

Kendr IntelligentAdaptiveMeasured after routing. The response reports actual usage against the Kendr plan balance.
GPT-5.6 Sol$0.80Standard rate for this context size.
Claude Opus 4.8$0.75Standard token pricing; server-side tools may add usage charges.
Gemini 3.1 Pro$0.32Standard ≤200K prompt rate.

Why Kendr Intelligent has no precomputed dollar figure: one request may use a low-cost model or a frontier model, and retries or fallback can change the final route. The live catalog exposes default-route credit quotes with Kendr's fixed 5% markup; the response receipt records the settled charge.

How Kendr Intelligent works

Efficiency is a routing decision

Kendr Intelligent aims to spend capability only where the task needs it. These are system mechanics, not benchmark scores or a promise that every routed answer costs less.

01

Understand the task

Classify intent, required capabilities, context size, tool needs, output format, and risk before selecting a route.

02

Optimize the context

Remove irrelevant material and preserve instructions and evidence so the selected model receives a smaller, clearer request.

03

Filter eligible models

Exclude routes that do not meet context, modality, tool, policy, availability, or structured-output requirements.

04

Balance cost and capability

Choose between eligible routes using calibrated capability, expected latency, and the task's cost policy.

05

Use bounded fallback

If the selected route cannot complete the request, retry only through eligible routes without weakening the original capability and policy gates.

06

Return a receipt

Return the selected public alias, decision reason, versions, fallback outcome, usage, settled cost, and latency without exposing provider secrets.

Read correctly

What the numbers do—and do not—mean

A context window is capacity, a token price is a unit cost, and a benchmark is a test result. None of them alone predicts the total cost or quality of a real task.

Comparable facts

  • Published input and output token rates
  • Published context and maximum output limits
  • Supported modalities, tools, caching, and batch options
  • Whether model selection and an auditable routing receipt are built into the system

Workload-dependent outcomes

  • Tokens required to complete the same task
  • Latency under real provider load
  • Retry and tool-call frequency
  • Answer quality, acceptance rate, and total task cost

For an internal buying decision, run the same private evaluation set through every option and compare accepted answers per dollar.

Sources

Official specifications

Provider data was checked on July 18, 2026. Prices and preview availability can change; follow the linked source before making a purchasing decision.

  1. 01
    OpenAI — GPT-5.6 Sol model reference: $5 input, $0.50 cached input, $30 output per million tokens; 1.05M context; 128K maximum output; long-context multiplier above 272K input tokens.
  2. 02
    Anthropic — Claude Opus 4.8 model reference and pricing: $5 input and $25 output per million tokens; 1M context; 128K maximum output.
  3. 03
    Google — Gemini 3.1 Pro Preview model reference and pricing: 1,048,576 input tokens; 65,536 output tokens; tiered $2/$12 or $4/$18 pricing based on prompt size.
  4. 04
    Kendr — Developer guide and usage and billing: model aliases, Kendr Intelligent routing behavior, response usage, and shared plan balance.
FAQ

Common questions

The important distinction is between selecting a model yourself and asking Kendr to select a route for the task.

Is Kendr Intelligent a foundation model?

No. It is Kendr's routing mode. It evaluates task requirements, selects one eligible answer route, and can use bounded fallback. It does not invoke a second model to verify the answer.

Does Kendr Intelligent always cost less?

No. It is designed to improve task-level cost efficiency, but a frontier route or fallback can cost more than one low-cost fixed-model call. Compare accepted outcomes per dollar on your own workload.

Why not publish one Kendr Intelligent price per million tokens?

Because the route is not fixed. Cost depends on the selected model, optimized input size, output size, tool use, and retries. Kendr applies the same fixed 5% markup in every routing mode, and the receipt reports the settled charge.

Which option is best for predictable model behavior?

Select a fixed kc-* model alias when you require the same model family for every request. Use kendr-intelligent when task-level routing is more important than fixing the provider model in advance.

Choose a model—or choose a routing policy.

Use a fixed model when its behavior and price profile are already right for the workload. Use Kendr Intelligent when the workload changes from task to task and model selection, token optimization, bounded fallback, and audit receipts should happen inside the system.

Open Kendr