← Back to notes

無擷 · Research note

The Shape of Ability

Why AI makes some people faster and others slower

2,200 wordsv0.4

Keywordshuman–AI collaboration · ability differences · cognitive offloading · research notes

I kept coming across numbers that refused to belong to the same story. Some people worked 40 percent faster with AI. Some wrote better stories that also sounded more like everyone else’s. A group of programmers, all deeply familiar with their own projects, took 19 percent longer. Each study made sense on its own. Together, they made me wonder whether “Who is better?” was simply too blunt a question.

1. Four results that do not quite fit together

Start with writing. Noy and Zhang asked 453 college-educated professionals to complete several writing tasks. People with access to ChatGPT finished about 40 percent faster, and the quality of their work rose by about 18 percent. Those who started behind caught up the most. Without ChatGPT, the correlation between the first and second task scores was 0.49. With it, the correlation fell to 0.25 [1].

A field experiment at Boston Consulting Group found something similar. In a group of 758 consultants, those who began below the average improved by 43 percent on tasks within GPT-4’s strengths. Those who began above the average improved by 17 percent. The researchers also included a problem the model was not equipped to solve. On that problem, consultants who used AI were 19 percentage points less likely to get the right answer [2]. The task clearly mattered as much as the tool.

In a short-story experiment, AI-generated ideas helped writers with lower baseline creativity the most. Writers who were already highly creative changed little. Yet the finished stories also became more alike. In the one-idea and five-idea conditions, similarity increased by amounts equal to 10.7 percent and 8.9 percent of the similarity range in the control group [3]. Individual scores could rise while the group as a whole lost some variety.

Then there is programming. In early 2025, METR recruited 16 experienced open-source developers and gave them 246 real tasks in repositories they had worked on for years. These seemed like people who would know exactly how to use a new tool. With AI available, however, they took 19 percent longer. Before the experiment, they expected AI to make them 24 percent faster. Afterward, they still believed it had made them 20 percent faster [4]. The clock and the feeling of speed pointed in opposite directions.

I do not think one of these studies must be wrong. They involved different people, different work, and different generations of tools. What they leave behind is a better question: when a particular person meets a particular task and a particular tool, what makes the combination work?

2. The question may be larger than who is better

Many experiments begin with a baseline test and divide people into higher- and lower-scoring groups. The method is clear and useful. It can also make ability look like a single line.

Two people can both score 80 and arrive there in very different ways. One produces ideas easily but leaves a desk full of half-finished work. The other can organize and execute almost anything, yet struggles to find a fresh direction. Their totals are close. Their experience with the same tool may not be.

Generative AI is especially good at a handful of concrete jobs: sorting scattered material, rewriting sentences, making formatting consistent, sketching a first draft, and pushing something half-made toward usable form. If these are the steps that exhaust a person, the tool may fill a real gap. If these are already that person’s strengths, AI can add a new layer of prompting, fact-checking, and removing canned language. A baseline score alone will not tell us which outcome to expect.

That is all I mean here by the “shape” of ability. Alongside a person’s overall level, we can ask where that person moves easily and where the work tends to stall. The idea is rough, and human ability will not divide neatly into a few boxes. Still, it reminds us that the same score can hide very different people.

3. A small clue from research on ADHD

I first thought about this while reading work on ADHD and creativity. The comparison can offer a clue, but it cannot explain human–AI collaboration by itself.

White and Shah found that adults with ADHD performed better on tasks that asked for many unusual uses of an object, but worse on tasks that required several clues to converge on one answer. The first kind of task rewards ideas spreading outward. The second asks those possibilities to narrow. The authors connected part of the difference to inhibitory control [5].

A later meta-analysis combined 89 studies of ADHD, anxiety, depression, and everyday creativity. It covered 35,271 people and 261 effect sizes. The pooled correlation was only r = −.06, and its confidence interval crossed zero. Results varied substantially across studies and changed with the trait and the measure of creativity being used [6].

It would be easy to turn that −.06 into “ADHD has no relationship with creativity.” The number does not support such a clean claim. It averages three kinds of psychological traits and many different measures of creativity. Relationships that point in different directions can look close to zero once they are blended together.

The lesson I take is a plain one: when several abilities are pressed into one total, the most interesting differences may disappear inside the average.

4. The simplest version of the hunch

My first impulse was to subtract a score for execution and organization from a score for idea generation. The more I considered it, the less trustworthy it seemed. Both measurements would contain error, subtraction would pile those errors together, and human ability plainly contains more than two parts.

For now, the hunch is easier to say without an equation: a tool may help most when the step it handles well is the same step where a person usually gets stuck.

The reverse should also be possible. If the tool does not cover a weak spot, it may only place another set of controls beside an existing strength. People who know a task well may also inspect AI output more carefully. They can spend more time checking, repairing, and rebuilding an answer than they would have spent doing the work directly.

This story can account for some of the earlier findings after the fact. None of those studies tested it directly. Noy and Zhang, the BCG experiment, and METR did not map where each participant was strong or prone to stall. The METR developers’ experience does not prove that their abilities were unusually balanced. Familiarity with the task, skill with AI, and demanding quality standards could all explain why they slowed down.

A serious test would measure three things at once: the steps a person handles comfortably, the abilities the task requires, and the steps the tool can genuinely perform. It could then ask whether a closer fit among the three predicts a larger real-world gain.

5. Six other questions that follow

1. Claims about AI age quickly

METR carefully described its 19 percent slowdown as a snapshot of early-2025 tools in one setting. Follow-up data in 2026 began to lean another way, but participant selection and timekeeping had changed as well. The researchers declined to announce a clean new estimate [7].

That does not erase the first result. It reminds us that every experiment captures a combination: one generation of models, one interface, one group of people, and one kind of work. Change the model or the workflow and the people who benefit may change too. Any claim that “AI improves productivity” should travel with a date and a setting.

2. Smoothness is not always what is missing

Someone who already organizes and executes well, but has trouble escaping an old idea, may gain little from automatic formatting and instant summaries. A more useful tool might keep offering counterexamples, awkward interpretations, or the quiet question: “What else could be true?”

Most products try to remove friction. Some kinds of thinking need the right kind of resistance. Reaching an answer sooner and thinking it through carefully are not always the same achievement.

3. What happens tomorrow to the weakness patched today

People have always handed some mental work to tools. Research on cognitive offloading examines how we move memory, calculation, and search into paper, devices, and networks [8]. A meta-analysis of the “Google effect” included 22 papers and 30,889 participants. It found moderate effects in some categories, but results varied by region and measurement, and some confidence intervals crossed zero [9].

This does not prove that years of generative AI use will weaken a particular human skill. It does make a long-term question worth following. If a tool organizes our material every day, do we learn from seeing good examples, or do we simply practise the task less? Short-term performance and long-term change can move in different directions.

4. Feeling faster is not the same as doing better

The METR developers slowed down and continued to feel faster. In a survey of 319 knowledge workers, Lee and colleagues also found that people sometimes treated “less effort spent on the work” as if it meant “less effort needed for critical thinking” [10].

A fluid interface can create the feeling that the situation is under control. To judge whether a tool helps, subjective ease and finished quality should be measured separately. Along with asking whether the tool felt useful, we can track time, errors, later recall, and whether the person can still perform the task without it.

5. Everyone can improve while the work grows more alike

In the Doshi and Hauser story experiment, some writers earned higher scores and the stories became more similar. Both outcomes can be true. When many people draw help from the same kind of system, each piece may become more complete while the collection loses a few unusual directions.

Evaluations of AI-assisted creative work should therefore look beyond the average score for each person. They can also measure how much difference remains across the group. Individual gain and collective variety deserve separate columns.

6. Which step did we actually hand to AI?

In the Phaedrus, Socrates speaks through the story of Thamus and worries that writing will leave people with the appearance of wisdom while weakening their own memory [11]. Abacuses, calculators, and search engines later attracted versions of the same concern.

AI is not the first tool to move judgment beyond the head. Writing, spreadsheets, institutions, and search rankings have long shaped how people think. AI may be unusual because it can take over several connected steps at once: finding material, summarizing it, setting priorities, suggesting a plan, and turning the result into polished prose.

Rather than asking only whether AI will replace people, I find it more useful to ask: Which step did I hand over this time? Did I inspect it? Where do the final judgment and responsibility now sit?

6. Where this idea could go wrong

First, words such as “ideas,” “organization,” and “judgment” sound cleaner than they are. A person who cannot begin a task may be dealing with working memory, attention, perfectionism, or a simple misunderstanding of the assignment. Giving the pieces better names does not make them easy to measure.

Second, the same person changes across tasks. A gifted storyteller may be a clumsy programmer. Someone who judges well in a familiar field may lean heavily on AI in an unfamiliar one. The “shape” may describe this person in this kind of work, not a permanent label attached to the person.

Third, experiments become harder to control as they approach real life. A forty-minute task is easy to compare but says little about a project that unfolds over several weeks. A long project is more realistic, yet it is difficult to give everyone work of equal difficulty. Any future study would have to live with that trade-off.

For now, I treat the shape of ability as a lens for looking at the problem. It becomes a useful explanation only if it can make clear predictions and survive data that refuse to support it.

7. Leaving it in the notebook for now

I selected these studies around one question, so this is not a systematic review. Nor is there yet a reliable way to measure the shape I have described.

Putting the studies side by side has still clarified one possibility. The difference AI makes may depend on more than who begins ahead. It may also depend on whether a person’s sticking point happens to meet a tool’s strength.

The next step is not to collect more examples that flatter the hunch. It is to design a test the hunch could fail. Only then will we know whether “shape” helps us see something real, or merely gives the uncertainty a better name.

References

1. Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187–192. https://doi.org/10.1126/science.adh2586

2. Dell’Acqua, F., McFowland, E., Mollick, E. R., et al. (2026). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. Organization Science. https://doi.org/10.1287/orsc.2025.21838

3. Doshi, A. R., & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28), eadn5290. https://doi.org/10.1126/sciadv.adn5290

4. Becker, J., Rush, N., Barnes, B., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/

5. White, H. A., & Shah, P. (2006). Uninhibited imaginations: Creativity in adults with Attention-Deficit/Hyperactivity Disorder. Personality and Individual Differences, 40(6), 1121–1131. https://doi.org/10.1016/j.paid.2005.11.007

6. Paek, S. H., Abdulla, A. M., & Cramond, B. (2016). A meta-analysis of the relationship between three common psychopathologies—ADHD, anxiety, and depression—and indicators of little-c creativity. Gifted Child Quarterly, 60(2), 117–133. https://doi.org/10.1177/0016986216630600

7. METR. (2026). We Are Changing Our Developer Productivity Experiment Design. https://metr.org/blog/2026-02-24-uplift-update/

8. Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688. https://doi.org/10.1016/j.tics.2016.07.002

9. Gong, C., & Yang, Y. (2024). Google effects on memory: A meta-analytical review of the media effects of intensive Internet search behavior. Frontiers in Public Health, 12, 1332030. https://doi.org/10.3389/fpubh.2024.1332030

10. Lee, H.-P., Sarkar, A., Tankelevitch, L., et al. (2025). The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. CHI ’25, Article 1121. https://doi.org/10.1145/3706598.3713778

11. Plato. Phaedrus, 274c–275b.

Cite this article

無擷. (2026). The Shape of Ability (Version 0.4). 無擷.

Permanent link:https://wuxie.ink/en/notes/the-shape-of-ability/