Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • It’s a spyware, but a sovereign one
  • What do you expect from running on "someone else computer?"
  • How so? I've opted out of using my data for training.
  • I can understand companies not wanting to leak sensitive data or information, so not wanting their inputs used for training.

    But the individual level moral perspective confuses me. You object to your own input being used for training "for free", but you're ok to use models which already slurped the data of millions of other people "for free"?

    I don't get it. Obviously this is going to be a controversial take on here (I am not blind to the sentiment), so, please help me understand. If you're a conscientious objector, why are you using the models in the first place?

  • There's something to this argument. I for one use an API proxy (Kagi Assistant) when I'm submitting anything remotely sensitive to an LLM. But for coding in my hobby project, I have marked "use my conversations for training" in the Claude app. I do it in the off chance that one day we will use stupid amounts of compute for personalized medicine and to fight cancer; code for scientific and technical computing is really complicated and only a tiny fraction of humans can produce it.
  • Private users also have sensitive information which, if leaked, could be used to their detriment (e.g training data getting hacked, etc).

    I don't see any distinction between companies and people when it comes to moral objections re training data - at least, if that's a thing, I'm unaware of it.

    I do agree with your last point re the hypocrisy.

  • Summary of the different scenarios

    Service: Vibe Plan: Non-Enterprise Default: Opted in Opt-out possible? Yes

    Service: Vibe Plan: Enterprise Default: Opted out Opt-out possible? Yes (admin-managed)

    Service: Mistral Studio/API Plan: Not specified Default: Not stated, I assume opted in Opt-out possible? Yes

  • There are so models that beats all of Mistral models, plus you can run many of them locally. Why would anyone run Mistral?
  • Sovereignty.
  • The OCR stuff is quite usable. Their models perform good at some very basic tasks, so I use them to diversify. But their current model lineup is really terrible.
  • Possibly because Mistral is a French-based company, so EU companies can cut the American umbilical cord a bit more.
  • Qwen 3.8 27B absolutely demolishes Mistral best offerings at coding, and you only need a 5090 or 2x3090 to run it.
  • Mistral OCR is really good and their small models are great for simple translations, I use them regularly.

    Of course I’m not always on the most bleeding edge forefront of the latest hyped model so there might be better alternatives, but I’m happy with what they provide

  • Sovereignty. Not everything is about performance.
  • To be honest, I have a hard time noticing differences with Claude (in Kiro) and Mistral Vibe, at least with what I use them for. They simply feel like talking to the exact same thing.
  • To have at least a choice of using a European trained model. Downloadable weights is not open source. Mistral is able to respond to any regulatory queries about training data and concerns.
  • For all the people here alleging that AI companies will train on your data even when you as a user explicitly opt out of training and their terms say they will respect that, etc - do you also believe that within a few years of them having harvested your data, you could perform "knowledge probing" on their models by prompting various questions that determine if they can near-verbatim reproduce your unique data, and then have enough other people do the same that you can then just launch a class-action lawsuit? Because if not... I've got a startup idea for you.
  • Submitted title was "Mistral now trains on user input by default, except on enterprise tier". THat's good information (if true) but best suited to a comment in the thread rather than the title (https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...).

    We've changed it to the article title now per https://news.ycombinator.com/newsguidelines.html.

    by dang
  • I could be mistaken, but wasn't Mistral openly championing themselves as a company and EU option that wouldn't do this/didn't do this?
  • The same mistral that has a patent for "code implemented tool calls"https://news.ycombinator.com/item?id=49243397#49243588 and said "Companies selling artificial intelligence models in Europe should pay a "levy" to support cultural industries".

    They seem to get a lot of slack just because they're European but every new article I see about them makes my opinion a little worse

  • You must have not been keeping up with the times :)

    Their big play this year was to write a "whitepaper" on the future state of EU economy, which is something that they'd like to hand of to EU leaders and part of that proposal was some kind of mandatory 10% sovereign AI spend, or some other nonsense like that.

    They are, at least, trying to make big enterprise (with tailored models, custom integration) and government policy plays.

    That just goes to show you how ineffective they are as well at making AI click as a usecase.

    And when they don't get ahead by their own terms they copy what they see ongoing with US AI labs. Le Chat, and Vibe.

  • Maybe they realized that they were falling behind too much? To me it seems like user-feedback on bad descisions by the AI once it's trained to a basic level is among the most important signals in tuning the model to perform better.
  • This is a hugely misleading editorialised title.

    The page title is "Can I opt out of my input or output data being used for training".

    Right at the top of the page it says "In certain cases, your input and output data (such as conversations, documents, and other user-provided content) may be included in Mistral’s model training programs. You retain full control over this processing and have the right to opt out of these programs at any time."

  • Agreed. I read the title as they would start using my data for training and I couldn't opt-out. After looking at the page and checking my app (I have Pro subscription) it seems like I can opt-out and my initial opt-out when I subscribed was preserved.
  • This is the case for all the major AI providers. Training collection is on by default, but you can opt out.
  • "You can opt out at any time"

    ....

    "Of course we'll randomly turn that option off for you aka FB style and hope you don't notice. There is zero legal liability for us doing so, so why wouldn't we".

  • I pay for a subscription to Duck.ai mainly because I don't want to be constantly fighting my vendor to protect my privacy.

    Microsoft already did a rug pull on me and opted me in to training months after I signed up with Github Copilot. It exhausting and ultimately futile to monitor these companies.

    It's not guaranteed that Duck.ai will continue to uphold its promise of not training on your sessions — if the company gets bought by Microsoft, it's only a matter of time before the switch to "you can opt out at any time". But since privacy is Duck.ai's brand, it will be somewhat harder for them to hide what they're doing should they betray their customers.

    I also don't actually trust that Duck.ai sub-vendors OpenAI and Anthropic will uphold whatever contract they have with Duck.ai — the whole AI business model is built on lawless consumption of others work.

    We'll ultimately have to run our own models locally, because it's impractical to defend against untrustworthy AI vendors.

  • I also pay for duck.ai's service. They're my main "chat bot" thing, and while I had already been paying for their other services when I discovered I also had a duck.ai subscription, I choose to use them primarily because privacy is their whole brand. I only have two minor complainst: I want my conversations to sync between mobile and desktop (while being E2E), and I want a way to organize conversations into "projects" or grouped chats.
  • You have to be rather naive if you don't think these companies don't simply train on your prompts with or without your consent. They literally scrape everything - legal or not - and claim its fair use to train on, including straight piracy

    The idea that they'll steal from everyone except you is just wishful thinking

    by 20k
  • Pinky-promises, embarrassing. I'd point to tinfoil.sh. I'm not a shill for tinfoil, I haven't even used it or looked past its homepage really, but if we are able to legitimately secure privacy, then training data questions are moot (and the market for training data would probably shift/expose).