Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Do I understand it right that people now claim an ownership of the actual mathematical methods? What next, people patenting the letters and the numbers? And then the sounds? This starts going ridiculous.
  • dude, proper attribution is part of scholarship and academics. Not giving credit to the people whose work led up to your solution is malpractice.
  • When you want your unpublished work unpublished, don't store them on other peoples computers. Especially not on people their job it is to use data to create money.

    Would this data moved through a hack to ChatGPT, this would be another thing, but like this. No pity at all.

  • I was trying to patent our maintenance tracking algorithm, which produces guaranteed weight loss or gain within 2–4 weeks by producing accurate calorie and macro targets for people to follow; in our test, it beats GLP-1s like Ozempic, Tirzepatide, and Retatrutide in results.

    But later we found that algorithm and math cannot be patented.

  • If I understand correctly OpenAI cannot provide it. Because the models are essentially black boxes, especially this far after the fact, determining if this result built on training data based on conversations about the problem is impossible. So unless they can prove those conversations were never used for training then there’s no way to know.
  • This is exactly the issue.

    What the mathematicians could do is reveal whether they had the data-sharing opt-out on or not. But curiously, as far as I've seen, none of them will answer that question!

  • What? Can they just look if they fetched certain documents/conversations and feeded them into the training loop?
  • It's the 'conversations' of some mathematicians with OpenAI. So the question is: did anyone feed the conversation(s) back as training data.
    by xxs
  • Only under gross negligence would it be unprovable: Did you use a model whose training set included user data? Did the transcripts of any of the agents include a tool call whose result including user data?
  • It might be a lost cause regardless, because this is under the assumption that we can trust OpenAI to be honest about their own investigation, which is unlikely.
  • Math is excluded from copyright. So you can use any piece of math you ever heard from anyone and publish it, whatever the context, I think.
  • Looks like OpenAI's PR department woke up and put their main man heaney-555 to work on shaping public opinion.

    Dude makes up like 50% of the replies here.

  • Am I missing something obvious? Isn’t this just a simple DB query to see the state history of the “Data Controls” → “Improve model for everyone” toggle in the settings? Just report whether that was ever on and over what time period.
  • Yes you're missing several obvious things. Even saving the last changed date (nevermind every change date or what the change was) for every setting for every user would be earth crushingly wasteful. The value by itself isn't even worth including in backups.
  • Multiple OpenAI staff have publicly said they cannot do that as accessing specific user settings without their consent (or legal requirement) violates their internal privacy policy.

    However, the mathematicians could easily declare whether they had the toggle on or off. Yet curiously, they will not say!

  • If OpenAI remained a full nonprofit looking to build an "OPEN" AI for the benefit of all humanity (not only the US or a few shareholders), I would have been happy to share my code, my work, and even label their data... This said, I don't blame them. It's a difficult mission to remain a nonprofit and, at the same time, have the required capital investment to build AGI.

    I'm not criticizing them, but I hope this race towards the first-best result or AGI doesn't blind them to making good decisions such as not using their users' data without consent.

  • Same way for me.

    Imagine what a genuinely openness-focused organization of this sort could be. Even if we imagined a commercial half, we could imagine a foundation with mass-membership, perhaps with a membership fee equal to 1/2 the typical personal subscription and functioning to set the direction, elect the board, etc., and then a commercial half which might be rough, tricky, deceptive, making deals with anybody.

    I think I'd have been fine with the commercial half being a bit of a monster, as long as I'm part of the members and we decide what sort of board it gets and there's a clear "this is basically controlled by the public" and if I were part of a club of this sort, I would, like you absolutely fill up a directory with texts and computer programs and careful annotations to aid training.

    and they could have had it. It could have been easy to make an organization like this. I think you still can. An international AI club, the members vote on what sort of training material may be supplied and for what intents, create some committees to review quality, and then everyone starts making their little games and RL environments and annotated stories and programs that ordinary LLMs misunderstand, and then they get together and fine-tune something, and if that works well they then get some staff and better training infrastructure and end up with a commercial half.

  • Andrew Wiles gave 3 lectures, and only at the end of the last one he announced that he solved FLT. Imagine someone from the audience announced in between the second and the third lecture that they proved FLT (using his ideas, obviously).

    Why is it OK if openAI does it?

  • > Why is it OK if openAI does it?

    OpenAI has a lot of money, a lot of rich investors who want it to go to the moon, and an extensive propaganda operation.

  • This is a great analogy, and one I'm surprised more people haven't raised. Instead all your hear are comments about mathematicians being sour losers, they should have know the model terms of service etc...
  • It's a little like we've gone back to the problem of the customer becoming the product.

    If I were a mathematician I would not my unpublished work to go into the hands of a competitor.

    If I were a lawyer I wouldn't want private details of my defense to be made available to the prosecution. Anonymous or otherwise.

    I wouldn't want the plot to an unreleased book to be suggested to another author.

  • If I were a business, I wouldn't want my proprietary business information into a hands of a competitor. But CEOs are racing to do it, and (at best) relying in flimsy contractual guarantees they have no means to verify.
  • Here is some (adjacent, but relevant) context, which has been posted elsewhere but I think is worthwhile to mention again here:

    https://mathstodon.xyz/@tao/117237320796901560

    Especially in recent years the mathematics community has worked very much in good faith, and a lot of effort is spent trying to give appropriate credit for ideas. Even when ideas are discovered in parallel if it turns out that previous work contained the same essential idea it is by and far regarded as best practice to give priority in this case. The point, in good academic practice, is to maintain the health of the practice at large.

    As Terence Tao explains in that post, the goal of mathematics is not only to solve big problems. And, even if one were very single-mindedly focused on solving big problems, it is still (in the long term) better to maintain the health of the community at large so that problems which are out of reach at the moment may be in reach again in the future. Good academic practice is one part of this culture.

  • This is actually bigger than that. This is an "AI is eating the world" situation.

    The math community was one of the first to be affected because RLVR makes Math an easier target for AI. We witness AI eating the math community.

    Software community was also affected due to similar and other reasons. All other professional communities will face a similar challenge in the near future.