Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Terrifying to think that some techbro is out there right now concocting plans for an "AI background check" startup.
  • AI is going to be a lawyer's wet dream.

    Imagine the ads on TV: "Has AI lied about you? Your case could be worth millions. Call now!"

  • strange to talk about it in the future tense. it's here and yep, it's an object of fascination for the legal system
    by ghjv
  • Unlike medicine, AI isn't regulated, so lawyers won't have anything to go after.
  • And also (an incompetent and lazy) lawyer’s worst nightmare.

    At least once a week, there is another US court case where the judge absolutely rips apart an attorney for AI-generated briefs and statements featuring made-up citations and nonexistent cases. I am not even following the topic closely, and yet I just encounter at least once a week.

    Here are a couple most recent ones I spotted: Mezu v. Mezu (oct 29)[0], USA v. Glennie Antonio McGee (oct 10)[1].

    0. https://acrobat.adobe.com/id/urn:aaid:sc:US:a948060e-23ed-41...

    1. https://storage.courtlistener.com/recap/gov.uscourts.alsd.74...

  • "False or misleading answers from AI chatbots masquerading as facts still plague the industry and despite improvements there is no clear solution to the accuracy problem in sight."

    One potential solution to the accuracy problem is to turn facts into a marketplace. Make AIs deposit collateral for the facts they emit and have them lose the collateral and pay it to the user when it's found that statements they presented were false.

    AI would be standing behind its words by having something to lose, like humans. A facts marketplace would make facts easy to challenge and hard to get right.

    Working POC implementation of facts marketplace in my submissions.

  • I doubt that could ever work. It's trivial to get these models to output fabrications if that's what you want; just keep asking it for more details about a subject than it could reasonably have. This works because the models are absolutely terrible at saying "I don't know", and this might be a fundamental limitation of the tech. Then of course you have the mess of figuring out what the facts even are, there are many contested subjects our society cannot agree on, many of which don't lend themselves to scientific inquiry.
  • From her letter:

    > The consistent pattern of bias against conservative figures demonstrated by Google’s AI systems is even more alarming. Conservative leaders, candidates, and commentators are disproportionately targeted by false or disparaging content.

    That's a little rich given the current administration's relationship to the truth. The present power structure runs almost entirely on falsehoods and conspiracy theories.

  • It might be rich, but is it false? Are Google models more likely to defame conservatives than not?

    I think plausiblibly they might, through no fault of Google, if only because scandals involving conservatives might be statistically more likely.

  • Just a day ago I asked Gemini to search for Airbnb rooms in an area and give me a summarized list.

    It told me it can't and I could do it myself.

    I told it again.

    Again it told me it can't, but here's how I could do it myself.

    I told it it sucks and that ChatGPT etc. can do it for me.

    Then it went and I don't know, scrapped Airbnb or used a previous search it must have had, to pull up rooms with an Airbnb link to each.

    After using a bunch of products I now think a common option they all need to have is a toggle between "Monkey's Paw" mode: Do As I Say, vs a "Do What I Mean" mode.

    Basically where the user takes responsibility and where the AI does.

    If it can't do or isn't allowed to do something when in Monkey Paw mode then just stop with a single sentence. Don't go on a roundabout gaslighting trip.

  • "Knowing what they don't know" is one of the big unsolved problems in llm training. There's a finite number of facts. There's an infinite number of non-facts.

    If it could do what you're suggesting, it wouldn't need to do it in the first place.

  • At some point we have to be willing to call out, at a societal level, that LLMs have been fundamentally oversold. The response to "It made defamatory facts up" of "You're using it wrong" is only going to fly for so long.

    Yes, I understand that this was not the intended use. But at some point if a consumer product can be abused so badly and is so easy to use outside of its intended purposes, it's a problem for the business to solve and not for the consumer.

  • Businesses can't just wave a magic wand and make the models perfect. It's early days with many open questions. As these models are a net positive I think we should focus on mitigating the harms rather than some zero tolerance stance. We shouldn't allow the businesses to be neglectful, but I don't see evidence of that.
  • Maybe someone else actually made up the defamatory fact up, and it was just parroted.

    But fundamentally the reason ChatGPT became so popular as opposed to its incumbents like Google or Wikipedia, is that it dispensed with the idea of attributing quotes to sources. Even if 90% of the things it says can be attributed, it's by design that it can say novel stuff.

    The other side of the coin is that for things that are not novel, it attributes the quote to itself rather than sharing the credit with sources, which is what made the thing so popular in the first place, as if it were some kind of magic trick.

    These are obviously not fixable, but part of the design. I have the theory that the liabilities will be equivalent if not greater to the revenue recouped by OpenAI, but the liabilities will just take a lot longer to realize, considering not only the length of trials, but the length of case law and even new legislation to be created.

    In 10 years, Sama will be fighting to make the thing an NFP again and have the government bail it out of all the lawsuits that it will accrue.

    Maybe you can't just do things

  • Placing a model behind a “Use `curl` after generating an API key using `gcloud auth login` and accepting the terms of service” is probably a good idea. Anything but the largest models equipped with search to ground generation is going to hallucinate at a high enough rate that a rando can see it.

    You need to gate away useful technology from the normies, usually. E.g. kickstarter used to have a problem where normies would think they were pre-ordering a finished product and so they had to pivot to being primarily a pre-order site.

    Anything that is actually experimental and has less than very high performance needs to be gated away from the normies.

  • Google sat on this technology for years and didn't release their early chatbots to the public for this reason. The problem is that OpenAI opened Pandora's box and recklessly leaned into it.
  • LLMs have serious problems with accuracy, so this story is entirety believable - we've all seen LLMs fabricate far more outlandish stuff.

    Unfortunately, it's also worth pointing out that neither Marsha Blackburn nor Robby Starbuck are reliable narrators historically; nor are they even impartial actors in this particular story.

    Blackburn has a long history of fighting to regulate Internet speech in order to force them to push ideological content (her words, not mine), so it's not surprising to see that this story originated as part of an unrelated lawsuit over First Amendment rights on the Internet and that Blackburn's response to it is to call for it all to be shut down until it can be regulated according to her partisan agenda (again, her words, not mine) - something which she has already pushed for via legislation that she has coauthored.

  • This is about Gemma, Google's open weights model. And specifically availability through AI studio. I don't think they'll make the weights unavailable.
  • There should probably be a little more effort towards making small models that don't just make things up when asked a factual question. All of us who have played with small models know there's just not as much room for factual info, they are middle schoolers who just write anything. Completely fabricated references are clearly an ongoing weakness, and easy to validate.
  • How do you know how much effort they're putting in? If they're making stuff up then they're not useful, I think the labs want their models to be useful.
  • Given that the current hypewave is already going on for a couple years, I think it's plausible to assume that there really are fundamental limitations with LLMs on these problems. More compute didn't solve it as promised, so my bets are on "LLMs will never not do hallucinations"
  • I don't think there is any math showing that it's the models size that limits "fact" storage, to the extent these models store facts. And model size definitely does not change the fact that all LLMs will write things based on "how" they are trained, not on how much training data they have. Big models will produce nonsense just as readily as small models.

    To fix that properly we likely need training objective functions that incorporate some notion of correctness of information. But that's easier said than done.

  • If you disable making things up, LLMs will not work. Making stuff up is literally how they work.
  • > effort towards making small models that don't just make things up

    But all of their output it literally "made up". If they didn't make things up, they wouldn't have a chat interface. Making things up is quite literally the core of this technology. If you want a query engine that doesn't make things up, use some sort of SQL.

  • LLMs by definition do not make facts. You will never be able to eliminate hallucinations. It's practically impossible.

    Big tech created a problem for themselves by allowing people to believe the things their products generate using LLMs are facts.

    We are only reaching the obvious conclusion of where this leads.

  • One of the things that really has me worried at the moment are people asking chatbots who to vote for ahead of upcoming elections.

    Especially in parliamentary democracies where people already take political quizzes to make sense of all the parties and candidates on the ballot.

  • If you're asking a machine which human you should vote for, you probably shouldn't be voting.
  • > One of the things that really has me worried at the moment are people asking chatbots who to vote for ahead of upcoming elections.

    Why ? Don't worry, everything will be fine. Sincerely, FAANG