Dedicated model · Available as managed deployment
Mistral Small 3.2 — a compact European-family workhorse with a 32,768-token context window — validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. An OpenAI-compatible endpoint on your own machine, operated by AxForge in Málaga, Spain.
Why AxForge
| Compact European-family workhorse | Mistral Small 3.2 with a 32,768-token context window, served under the model name mistral-small on your own OpenAI-compatible /v1. |
|---|---|
| Your machine, your endpoint | A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory) reserved for you, serving only your traffic — validated on AxForge hardware and operated by AxForge. |
| EU-hosted, zero prompt retention | Hosted in Málaga, Spain (eu-es-1). Prompts and completions are processed in memory — not logged, not retained, never used to train — the same policy as the serverless API. |
Specifications
| Model | Mistral Small 3.2 — open model, Mistral family |
|---|---|
| Served model name | mistral-small |
| Context window | 32,768 tokens |
| Hardware | NVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge |
| Rental term | Hour, week, month or year |
| Hardware pricing | €0.69/hour on demand · €0.66/hour by the week · €0.62/hour by the month · €0.55/hour by the year, excl. VAT |
| Managed service | Quoted per deployment |
| Region | Málaga, Spain (eu-es-1) |
Full details, benchmarks and FAQ on the Mistral Small 3.2 page. Prices exclude VAT.
How it works
| 1 | Request deployment — describe your traffic, context needs and rental term. |
|---|---|
| 2 | You receive the configuration, hardware rental and managed-service price in writing before anything is billed. |
| 3 | AxForge deploys Mistral Small 3.2 on a dedicated DGX Spark reserved for you. |
| 4 | Point your OpenAI SDK at your own endpoint with model mistral-small. |
| 5 | Adjust the term — hour, week, month or year — as your workload settles. |
Request deployment or sign in to start.
FAQ
Not on the serverless API — it is available as a managed deployment: validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. The serverless API serves Qwen3.8 27B.
32,768 tokens.
AxForge publishes only numbers it measures itself, and has not benchmarked this model on its nodes yet. For quality benchmarks, see the official model card.
A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory) rented by the hour, week, month or year, running Mistral Small 3.2 behind an OpenAI-compatible endpoint on your own machine, hosted in the EU with zero prompt retention.
Two parts: the DGX Spark hardware rental — by the hour, week, month or year, with longer terms earning the lower rate — and the managed service, quoted per deployment. Both are confirmed in writing before anything is billed.