Self-host / Europe / multimodal / code

Mistral AI

Medium 3.5 (Modified MIT) and Small 4 (Apache 2.0); EU regional endpoints; GPAI Code

Mistral Medium 3.5 (04/28/2026) is the 128B dense flagship under a Modified MIT license: 256k context, text and image input, configurable reasoning; it replaced Medium 3.1, Magistral and Devstral 2.

Mistral Small 4 (03/16/2026, Apache 2.0) is a 119B MoE with 6.5B active params, also 256k. Alongside it: Large 3 (41B/675B MoE) and Ministral 3 (3B/8B/14B), all Apache 2.0 and multimodal.

Since 08/11/2026 Regional Endpoints (Europe or US) are GA and a Priority Tier with SLA is in preview. Mistral signed the EU GPAI Code. Le Chat: Free, Pro $14.99, Team $24.99/user.

  • Verified2026-09-03
  • Models covered4
  • Sources6
  • Deployment channels4

Purchase decision (when to choose / when to avoid)

Choose if...

  • You want a European provider with in-EU processing (Regional Endpoints) and the option to move to self-host on the same models.
  • You build on your own infrastructure: Small 4, Large 3 and Ministral 3 are Apache 2.0.
  • You need good code, document and vision models at lower cost than top-3 (Small 4: $0.15/$0.60).

Avoid if...

  • You need a single, always 'best' frontier model without variant selection.
  • You lack MLOps competence for self-host - SaaS/API from top providers will be simpler.
  • You rely on older lines (Pixtral, Magistral, Devstral 2, Mixtral) - they are retired from the API.

Cost in practice (scenarios)

SMB (copy + documents)

Le Chat Pro $14.99/mo (includes $15 API credits) or Team $24.99/user/mo; API Small 4 $0.15/$0.60 per 1M tok.

  • moderate volume
  • EU processing via Regional Endpoint
API - code and agents (Medium 3.5)

$1.50 input / $7.50 output per 1M tok.; batch -50%, cache up to 90% cheaper on input.

  • code, reasoning and vision in one model
  • Priority Tier with SLA in preview
Self-host

Small 4 or Ministral 3 on your own GPU; Medium 3.5 (128B dense) needs multiple GPUs. Cost is infrastructure + maintenance.

  • MLOps
  • monitoring and guardrails
These are estimates/scenarios (not an invoice). Actual cost depends on context length, number of users, limits and retention policies.

Deployment / data / enterprise

Deployment channels

  • Mistral API (Regional Endpoints: Europe or US)
  • Le Chat (Free, Pro, Team, Enterprise: SaaS, customer VPC or self-hosted)
  • Self-host weights: Small 4, Large 3, Ministral 3 (Apache 2.0), Medium 3.5 (Modified MIT)
  • Mistral models at cloud partners (depending on region and offering)

Data policy

Training on data
API and paid plans: no training on customer data by default; check terms for the Free plan.
Retention
Depends on plan/service; Enterprise with configurable retention.
Data residency
Regional Endpoints (GA since 08/11/2026): choose processing in Europe or the US; Le Chat Enterprise also in customer VPC or self-hosted.
With self-host, data policy is on your side.

Enterprise readiness

Admin
Team: domain verification, data export, 30 GB per user; Enterprise: custom models, agents, white label.
SSO/SCIM
Enterprise: SAML SSO.
Audit
Le Chat Enterprise: audit logging and connectors.
DPA
Depends on agreement.
Certifications
Depends on agreement.
Good compromise: EU + ability to go self-host when scale grows.

Best use cases

  • self-host and on-prem (Small 4, Large 3, Ministral 3 under Apache 2.0; Medium 3.5 under Modified MIT) with full data control
  • EU projects with emphasis on GPAI compliance, model documentation and in-Europe processing (Regional Endpoints)
  • code and agents (Medium 3.5, Small 4, Codestral 25.08), documents and images (multimodal models, OCR 4.1).

Strengths

  • All main text models ship weights: Small 4, Large 3, Ministral 3 (Apache 2.0), Medium 3.5 (Modified MIT); integration with vLLM and TensorRT-LLM.
  • Medium 3.5 (128B dense, 256k) combines code, reasoning and vision in one model; Small 4 (119B/6.5B MoE) is the low-cost production line.
  • EU/US Regional Endpoints (GA since 08/11/2026), Priority Tier with SLA in preview, GPAI Code signatory.

Weaknesses / risks

  • Fast retirement of older lines (Pixtral, Magistral, Devstral 2, Mixtral) means watching API deprecations.
  • Self-hosting Medium 3.5 needs multiple GPUs (128B dense); locally Small 4 and Ministral 3 are the practical picks.
  • Safety and guardrails with self-host are the deployer's responsibility; evaluation documentation thinner than top-3.

Current models (examples)

  • Mistral Medium 3.5 (128B dense, Modified MIT, 256k) - flagship for code, agents and vision; API plus weights on Hugging Face.
  • Mistral Small 4 (119B MoE, 6.5B active, Apache 2.0, 256k) - hybrid instruct/reasoning/code; replaced Small 3.x, Magistral Small and Devstral Small.
  • Mistral Large 3 (41B/675B MoE) and Ministral 3 (3B/8B/14B) - Apache 2.0, multimodal, 256k.
  • Premier (no weights): Codestral 25.08, OCR 4.1, Voxtral Transcribe 2, Mistral Embed; open: Shieldstral 1.0 (3B, Apache 2.0).

Alternatives (if this model doesn't fit)