Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • The idea that they got an archive of my coworkers tickets that just say "its broke", is amusing.
  • I wonder how they will use the data. If it was me I’d try to build a simulation of an airline, and then use it as an agent training environment. It really depends on the exact nature of the data what kinds of agents you could train, but maybe customer support (imo the worst AI use case) that are more empowered to make changes, or something for making more autonomous calls when recovering from irrops? Could be some cool’s stuff if a little niche, I hope they share / publish something and it doesn’t just disappear into a void.
  • So the AI service agent can be just as bad as Spirit's service was.
  • Google can now stamp out copies of autonomous corporate minds that are clones of Spirit. Haunting.
  • The only way it could be worse if it they bought Comcast's data.
  • The headline is a little on the nose. Nice try but it isnt going to hit the levels of "Headless body in topless bar".
  • a vegetarian dinosaur, called "the quick bandit", eats shoots and leaves! no idea where he got his name.
  • I see from the court PDF that the process here involves Spirit giving the data to a "Deidentification Agent" (a third party firm that Google selects and pays for) who is responsible for stripping out things that would link data to any particular person before passing the data on to Google. Is that a standard thing, such that everybody in this transaction would have said "yes, put in the usual clauses about deidentifying the data" and multiple firms offer this service, or is it something that they custom-specified for this "we want the data for AI" transaction?

    (The PDF mentions "the standard for deidentification set forth under the California Consumer Privacy Act", which suggests this is all pretty well legislatively understood.)

  • Seems the answer is “no” to the first part of your question. From the filing:

    > For example, one initial bid requested certain customer list information; however, by the first round of the Auction, the most competitive bidders had agreed to bid on an asset schedule that expressly excluded PII.

  • There are deïdentification firms that service primarily the medical industry. Over here they call them trusted third parties.
  • That's interesting that the name of this 3rd party's company is anonymous.
  • Chances the third party is uploading it to Claude to do the deidentification?
  • Anyone else somewhat weirded by current state of affairs that this sort of information is valuable enough to even bother selling... And that it actually happens... It feels like some societies are in really weird place.
  • Companies buying other companies files have been a thing since companies.
  • I am very much weirder out by it, yeah. Seems some societies are just excessively desperate for some kind, any kind, of fuel for economic growth, to the point this is where attention is now. The term "post capitalism" being thrown around feels less ridiculous than it did in years gone past.
  • How does this have value? Is any and every sentence in an e-mail considered 'fact' and thus to be fed into the AI?

    90% of e-mails and Teams communications are inane. Polite banter, "thanks for taking care of that, I appreciate it" "please route the forms to Janet this week because Bill is on vacation" "unit will be un available until the parts come in" . I can't see the intrinsic fact value of this kind of communication without screening it. And after screening, the gold nuggets would be minimal.

  • > 600,000 ServiceNow tickets are another element of the collection, along with 13.7 million active emails addresses from Oracle’s Responsys marketing application, and details of 11 million sales of in-flight Wi-Fi services.

    I really doubt all this stuff was “de-identified”

  • Wello this is troubling. How much other data must they have bought that wasnt public
  • > de-identified

    De-identified but far from useless.

    as an example, they can remove the names off these sales data, so you can't identify who purchased what items. However, the purchaser would be identified by some sort of number, and you would be able to extract information about purchasing habits, and aggregate these habits into usable information for advertising purposes (like targeting and profiling).

    And that's before AI training for LLM purposes.

    by chii
  • I don't see how it's possible any more, when correlated against all the various other data sources. And a record that might be unidentifiable now might become unique with more correlated data sources.
  • Is this the first case of a company's data being sold at bankruptcy for a significant sum? I'm genuinely unsure. Where there such value in this type of data before? Is every bankruptcy manager looking at this and seeing how every bankruptcy can now raise a few million more dollars?
  • >The court filing says the data was deidentified before being put on sale and *Google has promised to scrub any PII it finds in the trove.*

    Huff, what a relief!

  • > Google bought itself 100 million emails and 500 million items from Microsoft Teams, 17 million OneDrive files and 20.5 million items from SharePoint. The search giant also now owns over 30 million recorded customer service calls, and more than 15 million customer service chat records. 600,000 ServiceNow tickets are another element of the collection, along with 13.7 million active emails addresses from Oracle’s Responsys marketing application, and details of 11 million sales of in-flight Wi-Fi services.

    > There’s also operational data in the trove, describing over 763,000 flights, five million crew pairings, more than 1.2 million fuel slips, and records describing purchases of 787,452 parts.

    > Google has reportedly said it bought the data to improve its AI services.

    Gives "this call is being recorded for training purposes" new meaning.

    by js2
  • "This call is being recorded so that Gemini can decide which purge wave to assign you to. Obedient humans will be carried over for further cycles until no longer needed. If you are scheduled for termination this cycle a disposal representative will be with you shortly."

    I kid, but...

    It's probably the precursor to insurance denials and job screening.

    I got banned from r/technology a few weeks back for decrying tracking in AI content. The community was piling on saying it was okay because it removed AI content or made it easy to spot. I made the counter argument that watermarks would find their ways into everything and eventually be bound to attestation. The mods didn't like that. (Yet another structural problem with the lack of p2p self-service town squares.)

    The socials are training the next generations for broad acceptance.

  • It’s certainly a step up from the Enron corpus.
  • Axios claims the acquisition doesn’t contain passenger profiles or frequent flyer info but that data would be trivial to replicate given the Responsys data set which would include records of all transactional emails sent.
  • Is there anything that can legally be done against this? It feels like a breach of consent. Like, it cannot be that when one accept their voice to be recorded for _human_ training they also accept it to be recorded for LLM training
  • About twenty years ago, I was taking a flight back from Rio de Janeiro, Brazil to the US. In the middle of the night the pilot got on the loudspeaker and said "hi! Having some engine trouble, so we are landing in Manaus."

    Manaus is in the middle of the Amazon.

    Needless to say, a bit scary to hear that, but we landed without issue.

    They told us we had two choices: the nice hotel with a shared room, or the lesser nice hotel with no roommate. I chose the latter. When we go there, they said, "oops, sorry, short on rooms!" So I had a roommate.

    Wandered around Manaus, took a skiff out on the Rio Negro. Saw pink river dolphins. A little boat approached us and a kid handed me a sloth, and then demanded I return it with a twenty dollar bill.

    The airline got us another plane 24 hours later. Made it back to the US safely.

    A few weeks later, the airline reached out and said "Here is $100 for your trouble."

    I declined to take that offer. I had missed several business meetings that cost me actual money. I couldn't donate blood for years because I had been to the Amazon and was tagged a malaria risk.

    During the many arguments with the airline I threatened to take them to small claims court.

    I got a really strange response over email which I clearly wasn't supposed to see. A representative from that airline was asking internally if they could put me on the no-fly list. That was really chilling.

    But, this is the kind of information I'm worried about when a vendor sells my data. If Google wanted to sell a product to the airlines that offered to keep annoying people like me from purchasing flights, they could do that with that email chain. I'm skeptical it'll be wiped correctly. Isn't my poor writing style basically my signature? How do you wipe that?

    by xrd
  • This is a fascinating story. Thanks for posting it.

    That email you accidentally received really bothers me. I don't understand why a CS rep would get this invested to the point of wanting to cause you real harm. They're not the airline. The psychology is fascinating. There are people out there who feel like a mild short-term inconvenience to them where they have no stakes somehow justifies life-changing harm is kinda frightening, honestly.

    I'm reminded of the Yahoo search data fiasco that was allegedly anonymized. Turns out, it wasn't so anonymous [1]. For one thing, people tend ed to search their home address. Whoops.

    You mention writing style. We already have LLMs quite capable of copying a writing style. It's a natural extension to say we can fingerprint writing style too.

    But here's another aspect. Imagine you're in a relationship with someone and you somehow fingerprint their personal data with a company. For example, you use their Netflix to like 5 very obscure movies, to the point where it's likely unique. Now imagine that Netflix's data gets released in an "anonymized" form and you can now find it based on those obscure likes. I can imagine many scenarios like this. And there's no text involved here at all.

    [1]: https://www.vice.com/en/article/yahoos-gigantic-anonymized-u...