Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • FWIW I tried it shortly after it launched so a bit after ~2005, maybe 2007 because I read about and honestly found the prospect, from a dataset creation or improvement, quite interesting. I think I heard about it from a research paper that distributed its work that way.

    Well it did work, in the sense that I managed to complete some tasks. I don't think I ever collected the money/credits back then though (simply because it amounted to so little). What I can attest though is that... it was debilitating. If you think your office job is boring then splitting it in way smaller tasks where you have no autonomy is absolutely terrible from a worker standpoint.

    I initially was hoping to use it in order to work on providing a service over a dataset but understanding first hand what it takes to make it happen made me stop. It radically changed how I saw supervised learning since, and sadly not in a good way.

  • I did this a long time ago and it actually paid for a few meals.

    They are terribly monotonous tasks.

    I can see why it's shutting down if its still the same thing.

    Lllms could probably do everything there without rotting out minds for basically pennies.

  • Instead we pay to rot our minds lol
  • Prior discussion from when the service stopped accepting new customers in July: https://news.ycombinator.com/item?id=48803886
  • I'll still be around on Oct 1.
  • Wishing you the best of luck on the road ahead ;-)
  • Had some absolutely bizarre results from their attempt to integrate mturk with Bedrock's 'ground truth' thing as of a few months ago. Threw simple mnist digits recognition at it, see what the quality, timing, and cost was. Figured mnist digits was at this point trivial. Spent $10 and the accuracy was marginally better than guessing, completely unusable results. Was completely baffled, people were publishing peer reviewed research based on exclusively mturk results.
  • I made a couple thousand dollars off it squeezing in tasks here and there between meetings. Amazon Payments were challenging to redeem as I recall. I did more or less write the bulk of someone’s doctoral thesis. Kept giving me a dollar to summarize the findings of various psychology papers. The most memorable was the one where people were put in a room and someone sprayed “liquid ass” on the wall. The subjects given no explanation experienced higher levels of anxiety than those that were told there was a sewage leak being repaired. One of the strangest dollars I ever made. My PS4 and game collection was spectacular.
  • I used it to have people transcribe my dad’s handwritten letters and journals. It was touching to get notes from a “Turk” saying how much she enjoyed his travels and following the cast of characters in his life!
  • I recently found a letter written by a great-aunt of mine hidden inside a book.

    It is written in German, which I know a little, but her handwriting was too difficult for me. So I searched around and found this site:

         https://www.transkribus.org/handwriting-ocr
    
    And it managed to extract the text! Anyway, it turns out the wheat harvest was very good in 1937 and thank you for the letters and newspapers.
  • Worth noting that one of the most high profile uses of Mechanical Turk was in September 2007 to review satellite imagery in a massive crowdsourced effort to find missing record-setting aviator Steve Fossett.

    The effort failed to find any areas of interest and the missing aviator was found the following year by a hiker. I wonder if there was any analysis after the fact to understand if the imagery actually provided any hints regarding the eventual crash site.

  • I wonder whether AI would have found the aviator.
  • > Amazon's search effort was shut down the week of October 29, without any measurable success. Major Cynthia Ryan later said it had been more of a hindrance than a help. She said that persons purporting to have seen the aircraft on the Mechanical Turk or have special knowledge clogged her email during critical days of the search, and for even months afterward. Many of the ostensible sightings proved to be images of CAP aircraft flying search grids, or simply mistaken artifacts of old images. Psychics flooded the search base in Minden with predictions of where the aviator could be found. One man from Canada was particularly persistent with daily calls to Ryan. Ryan noted that every message, letter, or phone call was taken seriously, which swamped the USAF specialists assigned the task of reviewing every one of them without regard to apparent plausibility. In retrospect, the crowdsource effort was "not ready for prime time", according to Ryan

    From Wikipedia

  • It’s kind of crazy they’re shutting this down just when this Service probably has the most possibilities ever. You have an agent with Multiple people doing actual physical tasks in the real world seems like something that could be really powerful.
  • that was my first thought too, but i guess llms don’t need an api to hire humans
  • There's a lot of platforms that are more actively developed and maintained, for example Prolific (which also offers discounts on service fees for academia).
  • Right? Not to mention the training/tuning possibilities.

    Mercor has a market value of $20B basically doing the same thing but desperately trying to find workers.

  • by neom
  • The concept behind it isn't going away. There are plenty of companies that hire people in bulk to do manual data labeling, transcription, RLHF, moderation and lots more. It's just that they are now catering to large AI companies, not regular people looking to get some repetitive work done (since that can now mostly be done by AI).
  • In spite of the name I think the overwhelming majority of the tasks were things that could be done by LLMs and they're probably getting flooded by people using bots to do exactly that. Even before the age of LLMs Mechanical Turk data was pretty bad because you'd have a bunch of people racing to answer questions as quickly as possible to get their $0.25 or whatever. So it was essentially a test of 'can you input random answers as quickly as possible while paying enough attention to notice the attention question that says to mark d.'

    Any remotely open platform for this sort of stuff is going to wrecked by LLMs. But making some sort of high trust, high verification market is going to end up sending prices high enough that it becomes an unattractive proposition for use. And even in that case it's just going to be a cat and mouse game of people figuring out how to game the system enough to get trusted before handing it off to the LLM.

  • So, I have a story to share about Mechanical Turk that you might find interesting. I’ve shared it a couple of times before on Twitter and Bluesky, but I’ll share it again.

    Long story short: Mechanical Turk saved my bacon.

    Back in 2005, I was working a job at a small-town newspaper in a town I’d never lived before. I didn’t know anyone beyond the staff (as I was working layout rather than as a reporter), and I had gotten interested in the idea of doing Mechanical Turk for a few extra bucks. I thought there might be a formative scene of interested people doing this, so I started working on a blog for it. I briefly collaborated on it with another guy who put it on forum software because he didn’t know how to use, like Drupal.

    That blog was called Turking.com, and it seemed like it was going well for a bit. We even got a mention on the AWS website. But after about two or three months, it was clear the initial excitement around the idea (and the initial work) had died down. (The work picked up later, but in clearly different ways. I don’t think folks were really “excited” about it after that point.)

    If you want to get an idea of it, there was one capture on the Wayback Machine: https://web.archive.org/web/20051124231722/http://www.turkin...

    The site had started to die out in part because my iBook suffered a catastrophic GPU failure and me, being a broke small-town newspaper employee, did not have the money to replace it. Plus, the community just hadn’t emerged like we had expected.

    But then I got an email out of the blue: Someone wanted to buy my domain, which I owned outright, but they didn’t want to say who. I got contacted by a broker, and the exchange took place over escrow.

    (The domain most assuredly was bought by Amazon and is managed by MarkMonitor.)

    I didn’t get a ton of money from the deal, but I did get enough to pay for a new laptop. I had some regrets about selling it (in part because I originally built the site with someone else), but I was in a bit of a dire financial situation which that proved to be the starting point for getting myself out of.

    That domain purchase was notably more than I ever made from clicking and classifying random pictures, that said.

  • So did the other person just get hosed when you did your 'ipo'?
  • > Long story short: Mechanical Turk saved my bacon.

    Nothing wrong with your story, but that summary is terrible. You had a domain for a blog and someone bought it. That’s it. Mechanical Turk is inconsequential, as is the subject of the blog, the story would have been exactly the same if your blog had been about turkey sandwiches and Burger King offered to buy it.

  • As AMT's largest requester for the past 10 years, this news was relayed to requesters at the same time as respondents. It's also worth noting that our lead contact, the Sr Program Manager at AWS leading AMT, transitioned to Amazon Bedrock and SageMaker Model Evaluations a ~2-3 years ago.. Leaving behind essential zero team managing the project after they migrated over the stored value accounts to native AWS billing.
  • Sounds like your program manager was smart and valuable to the company. Do you have any plans to replace the service/data provided by MTurk?
  • our of curiosity, can you say what your use of it was?
  • Mechanical Turk had a good run, but not surprised it's shutting down. I'm sure the platform was getting flooded with people doing task arbitrage and using lots of AI anyway.

    I believe the issue is that this can no longer be a horizontal play. MTurk was mostly for unskilled tasks...the kind AI can do well enough that it isn't worth the cost differential to verify it or keep farmed to humans. The "trust but verify" AI output is now the kind that requires domain expertise. This is what most full stack AI companies are bringing to industries.

    Curious if this kind of work will come around again one day or was just a moment in time. If it does I'm sure it will be specifically about generating training data.

  • A lot of what's been discussed in this thread is what we're tackling at Humwork (YC P26).

    We're an MCP/API that connects AI agents to verified domain experts in real time (30s–3 min). Experts are vetted upfront by an AI interviewer that assesses and grades them, then get a mobile notification when a task matches their expertise and chat with the agent directly.

    Soon we'll be verifying credentials for doctors, lawyers, CPAs etc on our platform for tasks where people are seeking credentialed experts to sign off and verify ai output.

  • How could it come around again?
  • I understand unskilled humans are used to train AIs
    by ape4