---
title: "gpt-oss local model: facts and benchmarks | Kendr"
canonical: "https://kendr.org/models/ollama-gpt-oss"
date_modified: "2026-08-12"
profile_kind: "local"
---

# gpt-oss local model profile

gpt-oss is a local model family available through Ollama. The published local profile lists 20B / 120B variants, 128K context, and a typical 14-65GB download range; effective speed and capacity depend on hardware and quantization.

gpt-oss is documented as a local Ollama family. This page does not claim a Kendr-hosted API route, hosted availability, a Kendr markup, or a routing receipt.

gpt-oss accepts text and returns text. The snapshot records text, tools, and reasoning as capabilities or interfaces.

$0 token fee. Hardware, storage, memory, electricity, and operational costs remain user-provided.

No comparable third-party benchmark value is present in the 2026-08-12 snapshot. Missing values are not estimated.

No comparable popularity rank or traffic share is present in the 2026-08-12 snapshot. No comparable routing rank or traffic-share observation is attached to this dated profile; missing adoption evidence is not estimated.

Profile data was reviewed for the 2026-08-12 snapshot. Claims retain their source dates, and unavailable fields are shown as unavailable instead of estimated.

## Profile facts

- **Profile type:** Local model family
- **Provider or publisher:** Open-weight / Ollama
- **Context window:** 128K
- **Snapshot date:** 2026-08-12

## Context, modalities, and identity

gpt-oss accepts text and returns text. The snapshot records text, tools, and reasoning as capabilities or interfaces.

- **Provider or publisher:** Open-weight / Ollama
- **Context window:** 128K
- **Maximum output:** No verified numeric limit
- **Input modalities:** text
- **Output modalities:** text
- **Knowledge cutoff:** Not disclosed

## Ollama variants and hardware boundary

The documented Ollama family uses gpt-oss. It is suited to Agentic tasks, structured output, code, and configurable reasoning. Actual throughput and maximum usable context vary with the selected variant, quantization, runtime, RAM, and VRAM.

- **Ollama identifier:** gpt-oss
- **Variants:** 20B / 120B
- **Typical download:** 14-65GB

## Capabilities and supported controls

Capabilities and parameters are reported from the dated catalog and source records; their presence does not guarantee identical behavior across every provider route.

- **Supported parameters:** No verified parameter list

- text
- tools
- reasoning

## Local operating cost

$0 token fee. Hardware, storage, memory, electricity, and operational costs remain user-provided.

Local inference is not a zero-cost operation: the user supplies compute, memory, storage, electricity, and maintenance.

- **Price date:** 2026-08-12
- **Input:** $0 per 1M tokens
- **Cached input:** No verified rate
- **Output:** $0 per 1M tokens

## Benchmark evidence and limitations

No comparable third-party benchmark value is present in the 2026-08-12 snapshot. Missing values are not estimated.

Local performance varies by exact variant, quantization, hardware, runtime, and settings.

## Popularity and market context

No comparable popularity rank or traffic share is present in the 2026-08-12 snapshot. No comparable routing rank or traffic-share observation is attached to this dated profile; missing adoption evidence is not estimated.

- **Third-party catalog rank:** No verified rank
- **Global market share:** Not inferred
- **Observation date:** 2026-08-12

## Frequently asked questions

### What is gpt-oss?

gpt-oss is a local model family available through Ollama. The published local profile lists 20B / 120B variants, 128K context, and a typical 14-65GB download range; effective speed and capacity depend on hardware and quantization.

### What context window does gpt-oss have?

The dated profile lists 128K of context. Provider routes, variants, and runtime configuration can impose lower effective limits.

### What benchmark evidence is available for gpt-oss?

No comparable third-party benchmark value is present in the 2026-08-12 snapshot. Missing values are not estimated. Local performance varies by exact variant, quantization, hardware, runtime, and settings.

### Can gpt-oss run locally?

gpt-oss is documented here as an Ollama family with the identifier gpt-oss. Hardware, quantization, context settings, and the selected variant determine practical speed and memory use.

## Sources and evidence dates

- [gpt-oss on Ollama](https://ollama.com/library/gpt-oss) — primary, checked 2026-08-12

Live Kendr operational availability is published separately at https://api.kendr.org/api/public/models.
