Shieldstral preview
Shieldstral logo
Research Free open weights (Apache 2.0) — self-host only

Shieldstral

3B open-weights multimodal safety classifier that scores text and images against policies you write in plain English, on a single 16GB GPU.

Updated 2026-08-10

8
AI Score / 10
Visit Shieldstral
Quick answer

Yes. Shieldstral is free to use.

Listed pricing: Free open weights (Apache 2.0) — self-host only.

Pricing verified 2026-08-10 Is it free? Pricing Is it worth it?
Best for
Trust-and-safety or platform engineers already self-hosting an open model, who need text and image moderation inside their own infrastructure with policies they can rewrite without a fine-tune.

Overview

Shieldstral is Mistral's content-safety classifier: a 3-billion-parameter multimodal model that takes a piece of text or an image, plus a moderation policy written in ordinary English, and returns a verdict on whether the content violates that policy. It shipped on August 4, 2026 as Apache 2.0 open weights with an accompanying technical report, and Mistral puts the hardware floor at a single 16GB GPU. There is no hosted Shieldstral endpoint announced — you run it yourself, even though Mistral has sold a hosted text-moderation API on La Plateforme since 2024.

The policy-at-inference design is the interesting part, but it is no longer the rare part. OpenAI's gpt-oss-safeguard models landed the same idea under the same licence in late 2025, and Meta's Llama Guard 4 already covers text and images in one classifier. What Shieldstral actually contests is the size bracket: it claims the multimodal, bring-your-own-policy behaviour of those models at 3B parameters and a consumer-class card, where the closest policy-as-prompt open rival starts at 20B and Llama Guard 4 sits at 12B under a community licence rather than Apache 2.0. If that holds up, the deployment story changes — a guard model small enough to run inline next to your own model instead of on its own box.

The headline claim is that it matches or beats models up to seven times its size across text safety, refusal detection, policy adaptability and multimodal tests. Classification is the workload where small specialised models close the gap, so the shape is plausible, but every number is Mistral's, on a task mix Mistral assembled, and "policy adaptability" was introduced in the same document as the model built to win it. Six days after launch there is still no published head-to-head against gpt-oss-safeguard-20b, which is the comparison that would actually settle whether 3B is enough. My read: worth it if you already run your own inference and want user content to stay inside your process; skip it if you handle a few hundred messages a day, where a per-call moderation API beats paying rent on an idle GPU.

Is Shieldstral free?

Yes. Shieldstral is free to use.

What the free tier covers: The weights are Apache 2.0 with no fee, no seat limit and no usage cap, so the real cost is compute: a 16GB-class GPU running inline with your own model. No hosted Shieldstral SKU was announced at launch; check Mistral's La Plateforme pricing page for current rates in case a managed endpoint has since appeared.

Listed pricing: Free open weights (Apache 2.0) — self-host only.

Pricing verified . That is the date the plans were last re-checked against the vendor's own pages, not today's date. Confirm on the official site before you pay.

Shieldstral pricing

Open weights Free (Apache 2.0)

Full model weights, commercial use, modification and redistribution permitted, explicit patent grant, no user or request ceiling. Self-hosting only — you supply the GPU.

Pricing verified . That is the date the plans were last re-checked against the vendor's own pages, not today's date. Confirm on the official site before you pay.

Is Shieldstral worth it?

Worth it for Trust-and-safety or platform engineers already self-hosting an open model, who need text and image moderation inside their own infrastructure with policies they can rewrite without a fine-tune.

You can test that on the free tier before paying anything. The recorded trade-offs are listed below, and any one of them can settle the question on its own.

The 8/10 AI Score is an editorial read of published capability, price and shipping pace. Nobody here has hands-on hours with Shieldstral. How we verify.

Worth it if

The strengths recorded against this entry.

  • Apache 2.0 with an explicit patent grant — commercial use, modification and redistribution with nothing to sign, unlike Llama Guard 4's community licence
  • Smallest model in the multimodal-guard class: 3B and a stated 16GB floor, against 12B for Llama Guard 4 and 20B for the nearest policy-as-prompt open model
  • Moderation policy is supplied as text at inference time, so one deployment serves products with opposite definitions of a violation
  • Scores text and images in one model instead of requiring a separate vision moderation service
  • Self-hosted end to end, which is the practical answer for EU and regulated products where data residency rules out a US moderation API

Not worth it if

Any one of these blocks your use case.

  • Every published benchmark is Mistral's own, and no independent reproduction has appeared in the week since launch
  • No head-to-head against gpt-oss-safeguard-20b, so whether 3B costs you accuracy against the obvious rival is publicly unmeasured
  • No hosted endpoint announced: you own the GPU, the scaling and the on-call, which prices out low-volume products that a per-call API would serve for pennies
  • Runs on the critical path — screening both input and output adds two passes per turn — and the plain-text policy is an unversioned behavioural surface that can drift without a code review

What sets Shieldstral apart

  • 3B parameters on a stated 16GB GPU floor — roughly a seventh the size of gpt-oss-safeguard-20b, the closest open policy-as-prompt guard model
  • Handles images natively alongside text, where OpenAI's open safeguard models are text-only
  • Apache 2.0 with a patent grant, against the Llama community licence covering the nearest multimodal alternative
  • Policy-at-inference itself is no longer a differentiator — OpenAI and others shipped it in 2025; Shieldstral competes on size and licence, not on the idea

Key features

Policy-as-prompt classification

Moderation rules are supplied as plain-English text at inference time rather than baked into the weights as a fixed category list. One deployment can enforce different rules for different products, and redefining a violation is a config edit rather than a fine-tune or a vendor ticket.

Native text and image scoring

Screens uploaded and generated images through the same 3B model that reads chat messages, so there is no separate vision moderation path to maintain. OpenAI's open safeguard models are text-only, which is the gap this fills.

Apache 2.0 open weights

Commercial use, modification and redistribution are permitted with an explicit patent grant and no user-count ceiling — a materially freer licence than the Llama community terms the nearest multimodal guard model ships under.

Inline guardrail deployment

Mistral states a 16GB GPU floor, small enough to sit on the same host as the model it is guarding and screen both prompts and responses. User content and moderation decisions never leave your infrastructure, which is the whole argument for it over a hosted endpoint.

How it compares

Compare head-to-head

Comparison explorer Put Shieldstral up against any three tools Opens with Shieldstral already loaded. Add up to three more from the full index and read pricing, features, pros and cons in one table.

Related reading

Ready to try Shieldstral?

Head to the official site to start with Shieldstral — pricing and plans are listed above.

Visit Shieldstral
← More Research tools
Weekly issue

The 5 AI tools that mattered this week.

One email, Fridays. No spam, unsubscribe anytime.