Writing

Artificial Artificial Intelligence Just Gained a Third Artificial

Mechanical Turk is closing. I was a business development manager for it in 2008.

Amazon Mechanical Turk stopped taking new customers on July 30, 2026. The notice went up a month before that, and AWS said in its SageMaker developer guide that it continues to invest in security and availability for Mechanical Turk but does not plan to add features. If you were thinking about becoming a customer, you missed the window by about twenty years. I spent a year of my life trying to get enterprises through that window.

The pitch

MTurk launched publicly on November 2, 2005. The origin story, which was the first part of my pitch was like all Amazon Web Service offerings, an internal problem that we fixed and the solution was now available to the rest of the world. In this case, the internal problem was duplicate product pages. Someone had to look at two listings and decide if they were the same thing. A computer couldn't do it and the work didn't justify a hire, so Amazon built a mechanism to farm it out, then opened that mechanism to outside requesters. The unit of work was called a HIT, a Human Intelligence Task.

Bezos named the service after the eighteenth-century chess automaton with a man hidden in the cabinet, and he called it "artificial artificial intelligence," a phrase that was and still is odd to say out loud. In 2006 he keynoted the MIT Emerging Technologies Conference on the theme of "the hidden Amazon" and highlighted three services: Mechanical Turk, S3, and EC2. Two of those three became cloud computing.

I was a business development manager on MTurk from September 2008 to September 2009. My job was to move MTurk out of the hobbyist column and into production at large companies. Several AWS leaders told me at the time that this was Bezos "Pet Project" which was the only reason it was still around. Several people in leadership saw it as a drain on resources. I cannot document this misalignment. It was hallway talk.

I presented an exercise which teams could do during and after my pitch. The exercise was to put a whiteboard somewhere central. Every time you catch yourself doing a task you have done before, with little variation, write it on the board. Do that for a month. Then break each item into the smallest units that still stand alone, the way you would decompose a calculation into steps a machine can run without knowing the goal. The leaves of that tree are your HITs. People liked the exercise. Some of them ran it, and it worked. Those deals still died, just further downstream.

Three departments each held a veto

Here is what it took to turn a whiteboard item into production work. A business stakeholder had to name the repeated task and decide to expose it. Then somebody had to engineer the workflow, meaning the decomposition, the quality controls, the redundancy scheme. Then a developer had to build against the API, because there was no business-facing interface (that's another story).

Any one of those handoffs could stall, and the person who had to move first, the stakeholder, had the least evidence that the whole chain would pay off. Multiply three modest probabilities together and you have my pipeline.

The close-rate data is long gone, but I kept the target list. It names 22 industries and about 110 companies, everything from Google and Walmart down to the Patent Office and NASA. Search, retail, real estate, staffing, transcription, translation, government. Breadth was the strategy. We figured any organization with repeated judgment tasks was a prospect, and the list shows how wide we cast. Reading it now is also a tour of the 2008 internet. Cuil is on there. Zune. Xanga. Myspace. The platform I was representing outlived them. What I can't reconstruct is the funnel underneath: how many conversations each name produced, how many reached a workflow design, how many reached production. My honest recollection is that production deployments were rare. And there is one detail on the list I only see now, all these years later. Two of the columns, transcription and translation, name the work that would eventually fill the platform.

What was missing

I did not have the vocabulary for any of this at the time. In research for this article I found it in the diffusion of innovations literature (Rogers 2003, ch. 6), and the mapping is almost embarrassing. Trialability, first. A stakeholder could not test the idea without borrowing a developer. The cheapest possible experiment cost an engineering favor. We, as stakeholders, hoard engineering favors.

Observability, next, and this one an elephant in the room that I constantly tried to overcome. I developed case studies but they were obscure. I had nothing where a stakeholder could look at a peer company and picture themselves. No one feared falling behind because there was nothing visible to fall behind.

And relative advantage had to be taken on faith, since the stakeholder funded the decomposition before finding out whether the decomposition produced anything valuable.

Who was on the other side

We had a story about the workers, and I told it. They were doing this for fun. Gamification, a thing you did with the television on. Beer money, not rent money.

Then someone ran the numbers. I no longer remember when the study happened or how far it traveled inside the company. What stayed with me is the shape of the finding: a small share of workers completed most of the HITs, roughly your standard 80/20 split, and for that group this was full-time work.

So the gamification story held for the typical worker. But the volume came from the tail, and we had been reasoning from the median. The story was justified through the typical ways we try and ignore wage inequality: cost of living is lower where those workers are, family support structures are stronger there than in the United States. Some of that is true, as far as it goes. All of the justification came after the finding, and I never saw anyone go looking for evidence against it. Including me.

A study published in 2018, drawing on 3.8 million recorded tasks, put the median Turker wage near two dollars an hour, with about four percent earning above $7.25 (Hara et al. 2018).

There are two questions here I cannot answer. Whether the finding changed how the platform operated, I don't know. Whether it ever reached the enterprise pitch, I don't know either. It did not reach mine.

What the platform became

MTurk outlasted my pitch by seventeen years. The use case that carried it was one I never presented because there didn't seem to be much money in it at the time. The mic drop wouldn't drop loud enough.

Researchers and businesses uploaded transcription and captioning work, and an anonymous international workforce did it for small per-task fees. Labeled data. The platform I sold as operational efficiency turned into training infrastructure for machine learning.

Then models got good enough to do the tasks themselves, faster and cheaper than people. Academic researchers had already been drifting away, partly because AI bots were posing as human respondents in their studies. Think about that one for a second. The bots had learned from data that workers like these labeled. Bezos's phrase had grown a third artificial, machine output passing as human judgment on a platform built to sell human judgment as machine output.

Six weeks before the closure notice, Bezos went on CNBC and argued that AI will elevate people rather than replace them. His analogy was a worker handed a bulldozer after years of digging with a shovel. He was talking about software engineers and radiologists. Mechanical Turk does not appear anywhere in the transcript, and as far as I can find, he has never commented on the closure.

The dynamo

The failure I lived through has a documented shape, which was a bit of a relief to discover. Paul David's 1990 paper, "The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox," follows electrification through American factories. Electricity was available in the 1880s. The productivity gains showed up decades later, and economists spent a long time puzzled about the gap.

Factory owners kept their old power system, one big engine driving overhead shafts and belts to every machine on the floor, and simply swapped the steam engine for a dynamo. Unit drive, a motor on each machine, did not arrive until the 1920s (Devine 1983). And unit drive is what made electricity desired because it freed the floor plan from the geometry of the drive shaft. David's argument is that the delay was the time it took for redesign, workflow restructuring, and skill development.

A dynamo bolted to the old shaft layout. That is my stakeholder in 2008, asking me which job MTurk would replace.

What changed

The three-person chain is gone now. You describe a task in a chat window and get output in the same sitting. No developer stands between you and the experiment, no API either. Trialability went from near zero to trivial, and on my 2008 model of the problem, adoption should have gone vertical.

The Anthropic Economic Index tracks task-level usage of Claude. Across a sample of 3,000 unique work tasks, the top ten account for 24 percent of the set, up from 21 percent in January 2025. Concentration is rising.

I want to be careful here, because I have been wrong about this platform category before. The data comes from just one provider and it counts conversations not value. Rising concentration fits several stories, including a benign one where the top tasks are where the value sits. I read it as weak evidence that the range of things people point AI at is not widening very fast.

Where the work moved

The middle of the chain got automated but the two ends did not. I am noticing this more and more everywhere I look. Upstream is noticing. Someone still has to see that they keep doing the same thing over and over. That is the whiteboard. A model will decompose any process you hand it (better than we did in 2008 with a marker). But someone has to outline the process. You notice a repeated task because it has annoyed you, and it can only annoy you if you do the work yourself or if the person who is doing the work is annoying or late or both. Then the annoyance accumulates into a thought.

Downstream is verification. The output reads well whether or not it is right, which is a new kind of problem. Someone has to check, and checking takes domain knowledge to catch a plausible wrong answer. That cost scales with volume, and it lands on the same person who was supposed to be freed up by all this. Cheap in the middle. Expensive at both ends. That is the shape of it as best I can tell.

Why noticing might be hard

Fair warning: this section is just my speculation. The candidates from cognitive psychology are functional fixedness, the difficulty of seeing a familiar object serve an unfamiliar function, and the Einstellung effect, the habit of applying a known method after a better one exists (Duncker 1945; Luchins 1942). I know both through the secondary literature and have not read the primary sources. Status quo bias predicts similar behavior by a different route.

The version I believe is that a repeated task you perform competently is invisible the way a well-fitting shoe is invisible. It produces no signal. The whiteboard works because it converts an absence of friction into something you can point at. I don't put much weight on the cognitive account. The organizational account covers the same observations with fewer assumptions about what is happening inside anyone's head, and David's factory story is an organizational account.

There is also the awkward fact that this essay contains two cases of the same blindness. AWS users could not see their own repeated tasks. My AWS colleagues and I could not see who was on the other side of the API. I wish we could have put those both under one mechanism.

Two industries, one platform

MTurk was early to a specific arrangement: work sold in units too small to add up to a job, priced in a market with no floor. You can see the same arrangement in rideshare, and again in the annotation work behind current models. And the same platform produced the labeled data that trained the systems now replacing it, which is the kind of loop you would reject in fiction as too neat. Other pieces of the gig economy were forming at the same time, and I was not close enough to that world to rank them.

What would change my mind

Start with noticing. Point a model at organizational telemetry, the ticket queues and calendars and mail and event logs and screen recordings, and see what it finds. If it surfaces automation candidates no one had named, and those get built and survive six months, then noticing was a data-access problem all along and my central claim is wrong.

Process mining is already a partial counterexample, and I want to be upfront about it. Celonis and its competitors have discovered repeated process patterns from event logs for years, since before language models. If process mining plus language models closes the gap, my whiteboard was a workaround for missing instrumentation and nothing more.

Then there is the build step. If firms with in-house engineering adopt at much higher rates than comparable firms without it, controlling for size and sector, then chat interfaces did less than I claim they did.

Or maybe slack is the real mechanism. If successful adopters differ from stalled ones on protected time, leadership mandate, and tolerance for a temporary dip, and not on any capacity to notice anything, then the cognitive story is decoration.

Verification gets its own test. Name a domain and a threshold in advance. If automated checking catches what human review catches there, the expensive downstream end is gone and half my barbell with it.

The lag frame could be wrong too. Top-ten share moved from 21 to 24 percent. If concentration drops sharply within two years with no visible reorganization, then diffusion is capability-driven and the dynamo comparison fails.

And last, the whole analogy could be decorative. MTurk had cost, latency, and quality-variance problems that language models do not have. If quality variance is what killed my deals, then I am generalizing from a case that does not transfer. I sat in business development, which is precisely the seat that would have hidden that from me. I put meaningful probability here.

Works Cited

Amazon Web Services. "Announcing Amazon Mechanical Turk." November 2, 2005. aws.amazon.com/about-aws/whats-new/2005/11/02/announcing-amazon-mechanical-turk/

Anthropic. "Anthropic Economic Index." anthropic.com/economic-index

Barr, Jeff. "We Build Muck, So You Don't Have To." AWS News Blog, September 2006. aws.amazon.com/blogs/aws/we_build_muck_s

CNBC. "CNBC Exclusive: Transcript: Jeff Bezos Speaks with CNBC's Andrew Ross Sorkin on 'Squawk Box.'" May 20, 2026. cnbc.com/2026/05/20/cnbc-exclusive-transcript-jeff-bezos-speaks-with-cnbcs-andrew-ross-sorkin-on-squawk-box-today-.html

David, Paul A. "The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox." American Economic Review 80, no. 2 (May 1990): 355-361.

Devine, Warren, Jr. "From Shafts to Wires: Historical Perspective on Electrification." Journal of Economic History 43, no. 2 (June 1983): 347-372.

Duncker, Karl. "On Problem-Solving." Translated by L. S. Lees. Psychological Monographs 58, no. 5 (1945): i-113.

Hara, Kotaro, Abigail Adams, Kristy Milland, Saiph Savage, Chris Callison-Burch, and Jeffrey P. Bigham. "A Data-Driven Analysis of Workers' Earnings on Amazon Mechanical Turk." In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, Paper 449, 1-14. ACM, 2018. doi.org/10.1145/3173574.3174023

"Jeff Bezos Talks Up 'Hidden Amazon.'" CIO, 2006. cio.com/article/260425/business-process-management-jeff-bezos-talks-up-hidden-amazon.html

Luchins, Abraham S. "Mechanization in Problem Solving: The Effect of Einstellung." Psychological Monographs 54, no. 6 (1942): i-95.

Pontin, Jason. "Artificial Intelligence, With Help From the Humans." New York Times, March 25, 2007.

Rogers, Everett M. Diffusion of Innovations. 5th ed. New York: Free Press, 2003. Chapter 6, "Attributes of Innovations and Their Rate of Adoption."

"Untold History of AI: How Amazon's Mechanical Turkers Got Squeezed Inside the Machine." IEEE Spectrum, 2019.