Mixtral 8x7B Instruct v0.1
Find hardware that runs Mixtral 8x7B Instruct v0.1 on your budget
Reference: Model card on Hugging Face · Published benchmark scores
- Parameters
- 46.7 B
- Active per token
- 12.9 B
- Architecture
- MoE
- Modality
- text
- Max context
- 32768
- Released
- 2023-12-10
- License
- apache-2.0
Measurements
Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.
No records yet for this model. Know of a published benchmark on it? Submit the link.
What verified, single-source and estimated mean, and the same rows with every filter and sort in the benchmarks explorer.
Speed over time
Cheapest hardware by target speed
The lowest-priced entry that reaches each threshold, using the strongest evidence available. Estimated rows are marked.