The most quoted statistic about artificial intelligence and employment was calculated in 2013. It was produced before the transformer architecture was published, before any commercial language model existed, and it has never been checked against what subsequently happened.
Carl Benedikt Frey and Michael Osborne's The Future of Employment is where the number "47% of jobs" comes from. Thirteen years later it is still the figure that opens conference keynotes, ministry strategy documents and pitch decks across this region. Almost every use of it is wrong in three separate ways: it is not recent, it was not a prediction, and it was not about job losses.
What the paper actually did
The method is unusual and worth stating precisely, because almost nobody who cites the number knows it.
Frey and Osborne convened a workshop of machine-learning researchers at Oxford and asked them to hand-label 70 occupations as automatable or not automatable. Seventy. That hand-labelled set became training data for a Gaussian process classifier, which then extrapolated a probability of computerisation to all 702 occupations in the O*NET taxonomy, using three "engineering bottleneck" variables: perception and manipulation, creative intelligence, and social intelligence.
The paper's own sentence is that "about 47 percent of total US employment is at risk," and the time horizon it gives is "perhaps a decade or two."
Read that carefully. At risk is a category, not an outcome. The 47% is the share of employment sitting in occupations the classifier placed in a high-risk band. The paper does not forecast that those jobs disappear, does not model adoption cost, does not model regulation, and does not attach a confidence interval to the estimate. It was published as an Oxford Martin working paper in 2013 and reached a peer-reviewed journal, Technological Forecasting and Social Change, in 2017.
We have not found a published back-test by the authors. Thirteen years is longer than the shorter end of their own stated horizon, and the American labour market has not lost half its employment. Whether the forecast was right is, remarkably, still an open question that its originators have not addressed in print.
The criticism that has stood since 2016
The strongest published objection arrived three years after the paper and has never been answered.
Arntz, Gregory and Zierahn rebuilt the estimate for the OECD using PIAAC data, which records what individual workers actually do rather than what their job title implies. Frey and Osborne treat an occupation as a single indivisible unit: if the occupation is automatable, every worker in it is at risk. The OECD rebuild allows for the fact that two people with the same job title do measurably different tasks, and that tasks within a job can change while the job persists.
Same question, different unit of analysis. The answer came back at 9% of jobs highly automatable across 21 OECD countries, against 47%.
Neither number is a measurement. Both are projections, and the OECD version has its own weakness: PIAAC task data is self-reported, so some of the within-occupation variation it finds may be measurement error rather than real difference. But the gap between 9% and 47% is not a detail. It is the difference between a policy problem and a civilisational one, and it is produced entirely by a methodological choice that neither paper's headline number discloses.
The later frameworks did not resolve it so much as change the subject. Eloundou and colleagues, whose 2024 Science paper is the closest thing to a modern successor, estimate that around 80% of the US workforce could have at least 10% of their tasks affected by language models, and about 19% could see at least half their tasks affected. Their own paper states plainly that this is a proxy for potential economic impact and is agnostic between augmenting and displacing workers. It is an exposure measure. It says nothing about whether anyone loses a job.
What has actually been measured
Set the projections aside and look at the observed record. It is thinner than the discourse suggests, and it does not point one way.
| Study | What was observed | Measured effect |
|---|---|---|
| Humlum & Vestergaard, Danish payroll | 25,000 workers, 7,000 workplaces, 11 exposed occupations | Precise null on earnings and hours; effects larger than 2% ruled out |
| Hui, Reshef & Zhou, online freelancing | Affected occupations on a large freelance platform after ChatGPT | Jobs −2%, monthly earnings −5.2% |
| Brynjolfsson, Chandar & Chen, ADP payroll | US payroll records through June 2026, workers aged 22–25 | 19% below a kept-pace counterfactual in AI-exposed occupations |
The Danish study is the most rigorous null in the literature: linked administrative records, difference-in-differences, and a confidence band tight enough to rule out effects above 2% two years after adoption. Users reported saving 3% of their time. The nulls held for heavy users, for early adopters, and for early-career workers specifically.
The freelance platform result points the other way, and the reconciliation is institutional rather than technological. Platform work has no employment protection, no firm-specific capital and instant demand reallocation. Danish payroll employment has all three. What determines whether exposure turns into lost earnings is the labour market's structure, not the model's capability. That distinction matters enormously for a region where informal and platform work is a large share of youth employment.
The Stanford payroll finding is the strongest negative signal that exists, and its authors are the most careful people writing about it. They describe their results as "early, descriptive indicators — canaries in the coal mine — rather than causal estimates," and they note that some divergence between more- and less-exposed occupations predates ChatGPT. The headline has also moved across versions: 13%, then 15%, then 16%, now 19% on data through June 2026. If you cite it, cite the vintage.
Acemoglu's macro model gives the ceiling from the other direction. Feeding the measured task-level cost savings from the existing randomised trials through a task-based model bounds total factor productivity gains at no more than 0.66% over ten years — about 0.064% a year — and GDP at 0.93% to 1.16%. Discounting tasks that are hard to learn pushes the bounds down to 0.53% and 0.90%. His implied share of tasks affected within a decade is 4.6%.
Frey and Osborne say 47% of employment is at risk. Acemoglu says 4.6% of tasks get touched. They are not measuring the same thing, which is precisely the problem: both numbers travel under the same headline.
What this means for the region
Every framework above was built on US occupational data. Frey and Osborne classified O*NET occupations. Eloundou and colleagues scored O*NET tasks. The OECD rebuild covers 21 OECD countries, none of them Arab.
Two studies have tried to extend exposure measurement to lower-income labour markets, and both produce numbers far below the ones in circulation. The ILO's refined index puts total generative-AI exposure at 11% of employment in low-income countries against 34% in high-income ones. The World Bank, applying an exposure measure to labour force microdata from 25 low- and middle-income countries covering 3.5 billion people, finds high exposure for 12% of workers in low-income countries and 15% in lower-middle-income countries — and states in the paper that exposure "does not equate to job loss."
MENA straddles that line. The Gulf sits in the high-income bracket; Egypt, Jordan and Morocco sit in the lower-middle bracket. There is no single MENA exposure number, and anyone who quotes one has invented it.
The larger gap is more serious. We have found no measured labour-market effect of generative AI in any Arab country — no employment study, no wage study, no productivity trial in Jordan, Egypt, Saudi Arabia or the UAE. Every regional figure in circulation is either a labour statistic that predates the question or a projection derived from American occupation codes.
What a founder should do about it
Stop citing 47%. It signals that the deck was assembled from secondary coverage. If a workforce investor knows the paper, the citation costs you the room.
Name the framework and what it measures, in the sentence. "Exposure," "at risk," "affected" and "displaced" are four different claims and only the last is about anyone losing work. A slide that says "47% of jobs will be automated" misstates its own source.
Separate projections from measurements, visibly. The measured record is three studies that disagree, one of them a precise null. That is a more honest slide than a single large number, and it is the one that survives an investor who has read the papers.
If your product is reskilling, price the fact that the exposure number for your market does not exist. Nobody has measured AI's effect on any Arab labour market. That is not a gap you can fill with an American index; it is an opening for whoever measures it first.
The 47% has survived thirteen years because it is memorable and because nobody who repeats it has read the method. Underneath it is a workshop, a hand-labelled set of 70 jobs, and a classifier extrapolating from them. It deserves to be cited as what it is: the first serious attempt at the question, made before the technology it is now used to describe existed.




