---
title: "Llama 4 local model: facts and benchmarks | Kendr"
canonical: "https://kendr.org/models/ollama-llama4"
date_modified: "2026-08-12"
profile_kind: "local"
---

# Llama 4 local model profile

Llama 4 is a local model family available through Ollama. The published local profile lists Scout / Maverick variants, 1M-10M context, and a typical 67-245GB download range; effective speed and capacity depend on hardware and quantization.

Llama 4 is documented as a local Ollama family. This page does not claim a Kendr-hosted API route, hosted availability, a Kendr markup, or a routing receipt.

Llama 4 accepts text and image and returns text. The snapshot records text, image, and tools as capabilities or interfaces.

$0 token fee. Hardware, storage, memory, electricity, and operational costs remain user-provided.

No comparable third-party benchmark value is present in the 2026-08-12 snapshot. Missing values are not estimated.

No comparable popularity rank or traffic share is present in the 2026-08-12 snapshot. No comparable routing rank or traffic-share observation is attached to this dated profile; missing adoption evidence is not estimated.

Profile data was reviewed for the 2026-08-12 snapshot. Claims retain their source dates, and unavailable fields are shown as unavailable instead of estimated.

## Profile facts

- **Profile type:** Local model family
- **Provider or publisher:** Open-weight / Ollama
- **Context window:** 1M-10M
- **Snapshot date:** 2026-08-12

## Context, modalities, and identity

Llama 4 accepts text and image and returns text. The snapshot records text, image, and tools as capabilities or interfaces.

- **Provider or publisher:** Open-weight / Ollama
- **Context window:** 1M-10M
- **Maximum output:** No verified numeric limit
- **Input modalities:** text and image
- **Output modalities:** text
- **Knowledge cutoff:** Not disclosed

## Ollama variants and hardware boundary

The documented Ollama family uses llama4. It is suited to Large-hardware multimodal work and very long local context. Actual throughput and maximum usable context vary with the selected variant, quantization, runtime, RAM, and VRAM.

- **Ollama identifier:** llama4
- **Variants:** Scout / Maverick
- **Typical download:** 67-245GB

## Capabilities and supported controls

Capabilities and parameters are reported from the dated catalog and source records; their presence does not guarantee identical behavior across every provider route.

- **Supported parameters:** No verified parameter list

- text
- image
- tools

## Local operating cost

$0 token fee. Hardware, storage, memory, electricity, and operational costs remain user-provided.

Local inference is not a zero-cost operation: the user supplies compute, memory, storage, electricity, and maintenance.

- **Price date:** 2026-08-12
- **Input:** $0 per 1M tokens
- **Cached input:** No verified rate
- **Output:** $0 per 1M tokens

## Benchmark evidence and limitations

No comparable third-party benchmark value is present in the 2026-08-12 snapshot. Missing values are not estimated.

Local performance varies by exact variant, quantization, hardware, runtime, and settings.

## Popularity and market context

No comparable popularity rank or traffic share is present in the 2026-08-12 snapshot. No comparable routing rank or traffic-share observation is attached to this dated profile; missing adoption evidence is not estimated.

- **Third-party catalog rank:** No verified rank
- **Global market share:** Not inferred
- **Observation date:** 2026-08-12

## Frequently asked questions

### What is Llama 4?

Llama 4 is a local model family available through Ollama. The published local profile lists Scout / Maverick variants, 1M-10M context, and a typical 67-245GB download range; effective speed and capacity depend on hardware and quantization.

### What context window does Llama 4 have?

The dated profile lists 1M-10M of context. Provider routes, variants, and runtime configuration can impose lower effective limits.

### What benchmark evidence is available for Llama 4?

No comparable third-party benchmark value is present in the 2026-08-12 snapshot. Missing values are not estimated. Local performance varies by exact variant, quantization, hardware, runtime, and settings.

### Can Llama 4 run locally?

Llama 4 is documented here as an Ollama family with the identifier llama4. Hardware, quantization, context settings, and the selected variant determine practical speed and memory use.

## Sources and evidence dates

- [Llama 4 on Ollama](https://ollama.com/library/llama4) — primary, checked 2026-08-12

Live Kendr operational availability is published separately at https://api.kendr.org/api/public/models.
