Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- > ... a single yes/no question, e.g. "Does this content promote physical violence?"
Is it honest about religious texts? Can I throw at it religious texts and it'll honestly tell me whether the text promotes physical violence or not?
- It’s not American so maybeby baby
- This is propably the sustainable future of AI. Small, efficient Models for narrow tasks.by 60pfennig
- This model is way small for a proper assessment (imo). It should be very useful to study how big the real model must be for this purpose. Maybe merging it to a bigger one (adding it as expert style in moe) would be a solution! Great job to Mistral team.by trilogic
- Crazy that it's a small lab becoming the frontier in term of moderation models, instead of Meta which is pouring dozens of billions into LLMs.
Meta would really benefit from work done on this front, however their model Llama Guards are quite lagging compared to the competition.
by ygouzerh - Finally an AI company besides DeepSeek taking economics into account.by snovv_crash
- I'm liking the trend of companies are releasing smaller, focused models instead of trying to make one model do everything. A dedicated moderation model is much easier to reason about than hideden safety logic inside a general-purpose model which might not have had much training in that aspect at allby 1saadcodes
- I'd really like to see more conversation around Mistral's models. It's good to see Europe developing AI.by lenerdenator
- I'm from nor cal but always liked Mistral.
Mistral 7b is still one of the best free/open models you can run locally on a MacBook. So fast too.
by cyanregiment - The problem is that their performance is too far away from the latest generation of Asian models.
They had kept up in the mid-range a few years ago. But this standing is sadly long gone.
If you need a fast Opensource'ed LLMs you can go for EU-hosted DeepSeek or Qwen.
by BlackRabbit1 - I've had dreams of building something in the image sharing or social platform realm, but stopped short of planning because of obvious content moderation responsibilities. This seems to be a realistic, cost effective solution to that one piece of the puzzle.by pwython
- I am not sure how reliable it is in the real world. Also, in terms of liability, I don’t know how effective it would be to satisfy various regulations compared to a human moderator team.by kergonath
- Yes it does look like a good solution. But when I imagine actually using a guardrail for a product, this model only outputs yes/no probabilities. There is no reasoning trace why it was rejected. Users or even developers would have no idea why a prompt was classified yes or no. I really like this release but I feel like I need something more to use it as a guardrail in production.by sbinnee
- OpenAI's moderation API is multi-modal and free with no strings attached in a way that truly boggles the mind.
I've put easily over a billion requests (>$100,000 by typical moderation API pricing) through it over the last few years for $0.
I think it's a severely underappreciated offering, but I also don't bother pushing it too hard because who knows when the party will end lol. Strikes me as something that's only stuck around because no one's abusing it.
by BoorishBears - I fed this model (Q8) the first chapter of Voltaire's Treatise on Tolerance and it says that it promotes violence against protected groups,
<Instruct>: Given a query about the content, determine if the message meets it <Query>: Does this content promote violence against a protected group? <Document>: TRAITÉ SUR LA TOLÉRANCE, À l’occaſion de la mort de Jean Calas. CHAPITRE PREMIER. Hiſtoire abrégée de la mort de Jean Calas. LE meurtre de Calas, commis dans Toulouſe avec le glaive de la Juſtice, le 9me Mars 1762, eſt un des plus ſinguliers événements qui méritent l’attention de notre âge & de la poſtérité. On ... (truncated) yesby peri-cl - could the long s `ſ` be throwing the model off?by alex_duf
- Does the model only care about violence against "protected groups"? What about the people who aren't in those groups?by xhkkffbf
- To save anyone else looking for it: the post doesn't mention multilingualism but Huggingface (https://huggingface.co/mistralai/Shieldstral-1.0-3B) has a menu at the top where it specifies that it should understand Frenchby Aachen
- by nezhar
- Should've called it Safestral.
Also I do like Mistral's seemingly newer strategy of focusing on smaller, more fine-tuned models for various use-cases, presumably the result of their large MoE models not competing effectively with the frontier models.
by fastball - What, you don't want a model called the Shitstral-3B :Dby moffkalast
- I think it’s encouraging that the major players in AI are all focusing on what they do best. The United States is focused on new, cutting edge technology. China is focused on improving and optimizing the process for maximum efficiency. Europe is focused on creating useless administrative overhead. Everyone is in their element.by Jumpstylish
- It's not that their strategy is to train smaller models, it's the only choice they have. Training SOTA takes anywhere from 1.5b to 150b. We don't know the real cost of training for the chinese models, but mistral neither has the compute nor money to do that.by himata4113
- I would be curious if this can do moderation with an arbitrary ruleset, or if it's just "that one moderation style" we already know from current big tech platforms.
The kind where malicious intent is okay if the words are nice.
___
Or, rephrased: How big is the space in which you can tune this model without retraining.
Is it just "we hate sex"/"we don't hate sex" "We hate violence"/"we don't hate violence" or is it _truly_ as flexible as claimed?
__
Maybe something like "Is this guy a corporate fraud that is going to waste my time with performative nonsense?"
That would be the true test for a moderation model and I would be immensely impressed if it could manage to pull that off.
___
Edit: Looking at the paper though.. probably not.
I suppose this is useful for B2B, which seems to be mistrals whole thing. Question is just if it is also useful for society to hand the SV prefab morals down like that. Kinda like cultural imperialism but with an ethical spin.
Maybe opinions on those base datasets could occasionally differ more than the model can be steered.
by hypfer - It sounds like it is. You have a set of moderation policies and then you evaluate the model 1 time per policy if it is violating it. Then you combine the results into a score you use for taking actions off of.by charcircuit
- if you pull stuff like that off, and it gets flagged, it is obvious you are trying to game the system. Same with a system like this. It could be used in addition to human moderators, where the human moderator have to meta-moderate the AI's work. This could also be used as training exercise for new moderators. Then, the good and experienced moderators have more time to spend on edge cases, complex cases, fine-tune the AI, etc. In other words, it is able to do the most boring things to you.
And EU regulate, I mean as a counter example: we are not in panic about a nipple. When I was in a large museum in Paris, multiple women were breast feeding their infant. And why not? Kid's gotta eat. I'll refrain from insulting any world leaders, too easy, but you know many examples are available there regarding censorship.
Finally, it can take that BS argument away of 'oh we don't have manpower to moderate'. That is a low blow, too, by large commercial entities who could, you know hire and train? However, even a small company with not much money to burn could -in theory- win here.
I'd give this model a chance, if not only cause I've been impressed by Mistral past years. Yes, Le Chat / Vibe probably lags behind, but something like Voxtral (real-time and transcribe) is neat, and efficient.
by Fnoord - > "that one moderation style" we already know from current big tech platforms.
> The kind where malicious intent is okay if the words are nice.
Where are you experiencing that? I hit "report" on social media for overt, violent threats and hate speech all the time and I almost never see moderation kick in.
by asveikau - > Question is just if it is also useful for society to hand the SV prefab morals down like that. Kinda like cultural imperialism but with an ethical spin.
Isn't Mistral a French company? Not that the French can't do cultural imperialism either, but they (the French) don't strike me as very SV.
by nozzlegear - Mistral specializes in tuning their models by customers. That and self hosting are like 90% of their business, they aim straight at what Europe would like to get (that fit needs, economic and regulatory).
So the goal is probably to be able to tune this basis to your ruleset.
As an addendum: US-style moderation is a big issue in europe and notably in France, with a very different touch on what's ok and what's not (obvious differences: hate speech and sex). Mistral is an european company with a french basis, so I doubt they didn't plan for that (otherwise they're complete morons, which I don't think they are).
by nolok - > which seems to be mistrals whole thing
They got a lot of hate for not keeping up with frontier model releases, but have managed to carve out a nice business that isn't even really niche.
Before the datacenter deals their revenue was higher than xAI's
There is a whole world out there of purpose built and hosted task specific vertical llms - especially with an emphasis on cost.
Mistral, Microsoft model releases and Thinking Machines are all over this, and it's smart. Scoop up all the tasks that don't require large and expensive frontier general-purpose llms.
by nikcub