Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • No self respecting mathematician is going to willingly become a marketing tool for these companies solving neglected and irrelevant puzzles.
  • i predict this perspective will not age well at all
  • No it's not, there are tons of programs in mathematics where you go there, form teams, work a problem as a team that's likely to get a result, and then publish the results from all the groups in the conference proceedings. They are called Research Collaboration Workshops.

    Very consistent pattern from these tech companies in their mathematics press releases that shows a conspicuous lack of experience in the research math world.

  • Will BigAI support this with free access to lots of hardware loaded with frontier models?
  • > It will be the first hackathon ever devoted to research level mathematics.

    Well, that's pretty damned ignorant; I was attending William Stein's hackathons on the BSD conjecture and the Sage Math project nearly 2 decades ago.

  • Interesting, if only I still have energy to work on something 40 hours non-stop
  • Should I read this as the big labs trying to move maths forward? Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs? On yesterdays "An Alien Mind" post from openAI they openly said that maths is not a priority for them, so I personally know what to think...
  • Yeah it's a bit of a tricky situation. It's almost like a bribe in a sense.
  • Probably many mathematicians want answers to the questions from the page:

    > This AI advancement raises the following questions: (a) How much can AI speed up the process from ideation to peer-reviewed publication? (b) What is the role of a mathematician when AI can solve conjectures faster?

    and the big AI companies agreed to sponsor them to find out because it is good publicity for the companies.

    by tzs
  • > Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs?

    I am told AGI has been achieved. If so, shouldnt these systems be out and about on their own ? Looking at 1st proof submissions in batch 2 it is clear that fully autonomous AI systems have a long way to go.

    AI harnessing human labor with the incentive of 2M in free tokens is the way my skeptic eye sees it, or humans being duped as reverse-centaurs.

  • Obviously the latter - this fact is betrayed by how the page lists the S^6 complex structure result, which was released as a 100 page barely readable mess (in fact even this might be too charitable), as still "unverified". Clearly a situation labs would like to avoid for future claimed results.
  • We are a student-led initiative. Our sponsors don't pay us and don't have a say in our decisions. All of our funding goes toward our judges and participants.
  • Hi check zeta.pukapasoft.xyz

    am I eligible?

  • I've applied as a team, hope I get in.

    Recently I care about how to create harness for mathematics that uses up the complete reasoning ability of the model. I care about both capability and cost.

    Most generic harness we have now are not made for maximizing reasoning. I've tested agents like codex, and rarely the cost of reasoning tokens reaches more than 20%. Which is quite strange as math requires a lot of reasoning. So hackathons can be a good test bed.

  • recent caltech grad here! and know some of the organizers well

    caltech's cs department is very, very weak, and has struggled to recruit top people in the last few years, and the most recent AI faculty hires have had issues. this is very slowly changing but a lot of the motivation for htis was to create a way for students to get ml "recognition" and learn about ai since it cant be done through the school right now. really glad to see hn picked this up!

  • We aren't affliated with the CS department. We're a student-led initiative. Our goal is to promote responsible AI use in math.
  • Won't deny that this is an interesting idea, but I feel like waiting on the output of an LLM for 40 hours feels like it is completely antithetical to what makes classic Hackathons appealing / educative.

    More generally, I don't think the shape of a hackathon (intensely working for a short timespan) maps at all onto the way LLM Math progress has seemingly been made so far; AFAIK it mostly involves picking out something for the Model, then having it run for a week with sporadic correction / encouragement.

  • Are you sure it’s letting it run and not going back and forth interactively?
  • If the goal is to accomplish something then why limit yourself with available tools?

    I’m not a full on AI optimist but it is absolutely the most powerful tool in a host of applications. From a Hackathon perspective, obviously in the 90s it was much more unorganized, but the same ethos existed. Use all available tools to accomplish the goal/task, it’s where a lot of incredible learning came out of. The same will hopefully happen in scenarios like this one

  • Have you done any math hacking with sol/astra or fable? It’s more fun than using them for coding. The models are great at the monotony, like constructing a Gröbner-basis, etc. But they’re all still absolutely awful at coming up with new ideas, new proof methods, or new constructive forms. So you spend all your time on coming up with novel hypotheses yourself and handing off the rote work to an agent.

    It’s also quite fun to get instant results by finding isomorphisms into unfamiliar areas of mathematics that previously would’ve required some networking in order to build a collaborative relationship.

  • But it's not "waiting on the output of an LLM for 40 hours" any more than a regular hackathon is "waiting for my damn teammates to finish their part for 40 hours". From my experience using agentic coding for hackathons, the best teams are those that coordinate with the AI agents in relatively quick cadence, generally giving it small tasks and steering it often. Teams may want to run some long-running sessions too, especially closer to the deadline, but even then, they'd probably want to run and follow several sessions in parallel, and continuously inspect their work so that they have reasonable confidence that their main efforts will wrap up before the deadline. There is an art to it.
  • 1. IMO the hard part isn't prompting. It's selecting the problem and understanding the solution. It'd be especially exciting if a participant formulates their own conjecture, proves it with AI, then generalizes it to a new theory.

    2. We have talked to mathematicians and frontier lab employees. We think 40 hours is enough to produce interesting results.

  • Hey guys, I'm one of the organizers. AMA.

    - We are a team of undergrads at Caltech. We don't represent Caltech, any Caltech departments, or any of our sponsors.

    - We don't receive monetary compensation. All the funding raised goes toward paying our judges and participants.

    - Our goal is to promote responsible AI use. You can read more about our commitments here: https://mathathonchallenge.com/faq.html

  • Serious question: Are you sure solving abstract mathematical problems with AI is responsible use? You will likely put mathematicians out of jobs, and I doubt solving the Collatz conjecture is urgent or will save lives. It also robs a future Fields medalist of the pride of doing all by themselves.

    All things AI seems to assume that more and faster is better, but there is no justification of that assumption. As a biological counterexample, a tree grown quickly will likely not be as healthy or strong as one grown slowly.

  • Hey! Any indication on what area of mathematics theses questions are from?
  • Is this hackathon only for those with formal math backgrounds?

    I've seen a few instances of AI assisted advances math and cs this year that were _not_ published by authors with formal backgrounds in those fields (or even institutional affiliation). Which makes me wonder if they would have a place at the event.

  • Why not specify 2-3 problems? Won't you have people show up having already spent a bunch of time on their self chosen problem? Which sort of defeats the point of seeing what you can do in a short period of time?