Model
Llama-3.1-8B-Instruct
Compact open Llama model for lightweight chat, drafting, and self-hosting
3 providersreleased knowledge cutoff 2023-12open weights
$0.025
best input /M · Inference
$0.025
best output /M
128K
context window
4K
max output
Compare providers (sorted by listed input price)
| Provider | Input /M | Output /M | Cache read /M | Cache write /M | Context | Capabilities | Status | Updated |
|---|---|---|---|---|---|---|---|---|
| Nvidia meta/llama-3.1-8b-instruct | — | — | — | — | 16K | no reasoningtoolsno structuredno vision | Jan 1, 2025 | |
| Inference meta/llama-3.1-8b-instruct | $0.025 | $0.025 | — | — | 16K | no reasoningtoolsno structuredno vision | Jan 1, 2025 | |
| Merge Gateway meta/llama-3.1-8b-instruct | $0.22 | $0.22 | — | — | 128K | no reasoningtoolsno structuredno vision | Jul 23, 2024 |
Price badge
https://models.sutraworks.ai/badge/meta/llama-3.1-8b-instruct.svgLive SVG, regenerated on every hourly sync — paste it into a README to always show the current cheapest listed price. Updates when prices change; no build step needed on your side.
Weights
Prices are per million tokens (USD). “—” means the provider does not publicly list a price for this model. Data from models.dev; verify with the provider before purchasing.