Editor’s Note: IFP’s Matthew Esche helped lead the Economics of Ideas, Science, and Innovation series this year, and facilitated several sessions, including this one. He’ll be introducing sessions that he facilitated.
When I have heard people’s grand visions of how AI will reshape scientific progress, I often find myself thinking: a few more people should talk to an economist. I have no doubt that AI systems will be revolutionary for science. Earlier this week we got a glimpse of it in mathematics. But economists inject the right kind of realism into the conversation.
Thankfully, this week’s Economics of Ideas, Science, and Innovation lecture brings us one of my favorite economists on AI.
Kevin Bryan has spent time writing benchmarks for AI and is a professor of strategy at the University of Toronto.
Visions of a million Einsteins in a data center will take time to transform many economically relevant undertakings, especially in the sciences. Even for our own Einstein, it took about 100 years to detect gravitational waves directly. Science runs on decades of institutions built around systems, many physical, all of them working together: lab instruments, journals, grants, to name a few. For advances in AI to have the largest positive effect, current institutions will need to be rearranged around it, and new ones built around AI from scratch will likely form too.
Kevin provides a compelling example in the lecture. When factories switched away from steam engines to electric motors, productivity barely moved. Factories were built around a single shaft overhead with belts running down to each machine. To paraphrase Kevin:
Adding an electric motor to a factory that’s built around belts and shafts barely matters. Reorganizing your factory around electricity matters a lot.

Reorganizing firms isn’t the only model to consider for how advances in AI plug into our economy and our scientific enterprise. Kevin presents us with a handful of models of growth and productivity. None are exactly right, but in certain applications, some are more right than others. As we’re reminded in the lecture, “all theory is wrong; it’s for a specific use.” Yet thinking through each of them is an enriching exercise, and for that, I’ll leave it to Kevin.
Here are the readings that go with this lecture:
David, Paul A. “The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox.“ American Economic Review 80, no. 2 (1990): 355-361.
Agrawal, Ajay, Joshua S. Gans, and Avi Goldfarb. “Exploring the Impact of Artificial Intelligence: Prediction versus Judgment.” NBER Working Paper 24626 (2018).
Shahidi, Peyman, Gili Rusak, Benjamin S. Manning, Andrey Fradkin, and John J. Horton. “The Coasean Singularity? Demand, Supply, and Market Design with AI Agents.“ In The Economics of Transformative AI. University of Chicago Press (2025).
Patwardhan, Tejal, et al. “GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.“ arXiv (2025).
Ide, Enrique, and Eduard Talamàs. “Artificial Intelligence in the Knowledge Economy.” Journal of Political Economy (2025).
Thanks to William Higbie and Harry Fletcher-Wood for their support in producing this video and the transcript.
Lecture Transcript:
The Economics of AI and Innovation
This is a topic that we’ve had filtering into some similar programs we’ve run — I teach the PhD innovation course up here at Toronto as well, and we’re always incorporating more and more of this econ. We’re at the point now where it just has to be a core component of the course.
For those of you who don’t know me, I’m trained as a game theorist — a weird game theorist who studied under Joel Mokyr. I’ve always been interested in innovation, and I’ve been doing work in the theory, history, and empirics of innovation for many years. I started working in AI quite a while ago. I came to Toronto in 2014. If you know your history of AI, 2012 is the big ImageNet competition, where a team of students and professors from this university used deep learning to win a major contest for the first time. That led to a complete blow-up in Silicon Valley. When OpenAI starts, one of our PhD students here, a guy named Ilya Sutskever — Elon Musk and Google are having a bidding war to hire him, because even by 2015 it was widely understood in Silicon Valley where AI was going.
In 2017 we have the famous Go match — the move 37 match — where a deep learning model beats the world champion. Then in 2022 we have ChatGPT. A lot of people now, when they hear AI, they’re thinking of large language models you interact with in a chatbot. That won’t be all I’m talking about. AI is a much broader problem. I’ll give you a bit of structure on how we think about it. But in order to make sure we’re all on the same page, I need to show you a handful of charts, which are hopefully going to convince you that AI as of March 2026 is relatively uninteresting. The interesting thing is the lines on the graph and their implications.
This is the time-horizon benchmark from METR, a private organization.
They set up a small number of computational tasks, where they wanted to see how long a model could go until it screwed things up. How long does it go, 50% of the time I run the model, before it makes mistakes and destroys the task? A few years ago the answer was a minute or less. You can see on this graph, we get up to late 2024 and we’re talking a half hour. This is from January — it’s four and a half hours. Actually, the benchmark is broken now. They tried to measure a new model and got 14 hours, and realized the problem is it’s just too saturated to really know at this point — the variance has gotten too big. This benchmark’s already destroyed, that quickly.
Here’s another one that I like, called FrontierMath.
They hired top mathematicians to write problems they thought were hard. Tier 3 is: “Someone in my field could solve it, but it would be pretty difficult; for a PhD student this one would be quite a challenge.” As of November 2024, it’s 2% solved. By December 2025 we’re talking just over 30%. Now we’re up over 50%.
They had to make a new benchmark called FrontierMath Tier 4.
These problems are really, really hard — legitimately difficult problems written by people like Terence Tao out of their file drawers. This was meant to be one that was going to last a few years. They release it in late 2024. No model solves any of the problems until — I want to say February 2025 — o3-mini solves one. Then we start cooking, and we’re going up and up. Now we’re almost to 40% on FrontierMath Tier 4.
They had to create another new one they’re going to release, called FrontierMath Open Problems, which are literally unsolved problems in the file drawers of some of the world’s best mathematicians. On the math side, that’s where we’ve gotten to with these models. Forget about the absolute number we’re at now — just follow the line on the graph up. That’s the interesting thing.
We’ve gotten to the point now in 2026 where — I write benchmarks for one of the big labs. I’m very close to not being able to do it. I can’t come up with problems that AI consistently cannot answer. There are some classes where we know there are issues: putting together a thousand answers correctly in a row is still difficult; consistency is difficult. But these problems are all quite solvable.
Even the type of things that large language models are going to have trouble with — large language models know how to use tools through agentic frameworks, and large language models are at the point where, inside the big labs, they’re contributing to research on the next model. So even if the breakthroughs we need aren’t large language models — they’re other forms of a deep learning model — it doesn’t matter, because large language models now are sufficient to speed up the research on developing that next model.
If you want to see where this is heading — a guy named Leopold Aschenbrenner wrote ‘Situational Awareness’ in 2024. You may have known him if you’re an economist. Six or seven years ago he wrote this growth theory paper — it’s a pretty good paper. When it was put online, he was at Columbia, and people were like, “Who is this kid, some new grad student?” No, he was a 19-year-old undergraduate. So he decides not to do his PhD in econ. He goes to OpenAI, because apparently he made good decisions the whole time, writes this book, gets fired from OpenAI for — I don’t know — talking to the press too much, and then he starts a hedge fund for AI investments and now makes a bajillion percent return per year.
He wrote this long essay called ‘Situational Awareness’, where he tries to explain what people are thinking inside the labs. He’s going through some of the benchmarks they had been trying to make.
For example, we tried to make this MATH benchmark in 2022 — doesn’t matter, it’s already solved by 2024, there’s no point in looking at it more. Then we try to create GPQA. These are like PhD-level qualifying-exam questions across fields. You wouldn’t do well — an in-domain PhD would get 80% on average. And he goes, “I expect this to fall within a generation or two.” We’re 21 months later, and there are models well into the 90s on GPQA. It’s more than fallen.
The reason the lines on the graph matter is: if you look at some of the other graphs that Leo talks about, we’re going to project those types of increases in capability forward another two or three years, and then things start getting interesting/very weird.
Particularly for economics of innovation, they get interesting/really weird, because we start to get to the place where, at least on technical grounds, it seems like AI should be able to contribute to speeding up the rate of science in a really meaningful way.
When Dario Amodei at Anthropic talks about a million Einsteins inside a data center, he’s imagining that you’re going to get a huge increase in growth because of these technical capabilities. We’re going to go in a slightly different direction. We’re going to take the technical projections from AI experts seriously, but we’re going to look at the economic reasons why that outcome might or might not happen, and where there might be remaining unknowns or open research questions that matter for whether in 2030 we have a million Einsteins in a data center doing a million Einsteins’ worth of work and causing an appreciable rise in growth.
We don’t traditionally do big-lab or big-money research in economics, with a few exceptions like Raj Chetty’s group. We work on whatever we want to work on; we have control. There are some kinds of papers, like GDPval, that are not really possible to write in that format.
It does make me think whether there are different ways to set up the internal production of science that we might want to pursue in a world with a lot of AI.
Let’s put all this together. This is what we’re going to take for granted — you can disagree; I’m just going to say you’re wrong if you disagree.
The vast majority of tasks that you can do on paper or on a computer, AI will be able to do in the near future. Shane Legg, the co-founder of DeepMind, would call this minimal AGI. Past that, I don’t know — a lot of uncertainty. It might depend on robotics advances and on policy. But at the very least, a lot of important tasks are going to be doable by AI, or are doable already.
Despite that, the technical improvements in AI will not map very quickly into productivity or growth. To give you one example, that I got from Ben Jones, if it is 1960 and you need to run a regression, you’ve got to invert a matrix. Inverting a 10x10 matrix — I don’t know if you were ever forced to do this in your undergraduate math courses, but it takes forever. It’s a very simple matrix to invert. At this point the cost of inverting a matrix is whatever the smallest epsilon you want it to be.
Despite the fact that we are so productive on inverting matrices that I can invert as many as I want and not even think about the cost, it barely affects growth. The reason is that any additional improvements — if we made the cost of inverting a matrix 90% cheaper at this point — it’s the Baumol effect; so unimportant in the production function compared to other bottlenecks that further improvements basically don’t matter. At least for the kind of work we do. It might matter for computer chip efficiency, but it barely matters. The mapping between “a technology gets 10 times better” and “growth gets 10 times better” is not going to work. We’re going to see why.
The dynamo and the computer
Let’s start from our friend Paul David. I love this paper, ‘The Dynamo and the Computer.’ Everyone reads this paper. I think it’s come up when I’ve talked to private-sector folks half a dozen times — somehow it’s made the jump from academia; they’re all aware of this one.
This is a factory prior to electrification.
You can see on the top I’ve got a shaft. It’s going to turn these belts here; they’re going to hit the individual machine and provide power to the individual machine. Somewhere in the basement I’ve got a water wheel or a steam engine that’s creating power. I’ve got one source of power for the whole factory. So it’s either turned on or turned off — meaning I want all the machines to be used at the same time.
It’s somewhat dangerous. You probably notice those old garment-district-type buildings in New York City — tall brick buildings with heavy floors. That’s because of all the shaft apparatus we need up here in the sky and the heavy machines that we have to bring these belts down to. We’re really restricted in where we place the machines. We often are going to do them stacked vertically in order to move the power more efficiently. So the entire design of the factory is based on these belts.
It’s 1882. Our friend Thomas Edison founds General Electric and sets up the first power plant in New York City. We can turn on the lights now. We can just plug in the machines. How much efficiency are we going to get from this invention? Initially the answer is hardly any. If it cost us five cents of electricity to turn this machine all day, now it’s going to cost us four cents because we plug it in instead of using the steam engine. Who really cares? It barely matters for the overall productivity of the firm. If our costs were 5% power and 95% other things, we just saved 1% of costs. You’re barely going to notice this in productivity figures. But you probably have a suspicion that electricity matters a lot for productivity — and you’re right.
What Paul David points out is that this factory barely gains from electricity. This factory, in River Rouge in Detroit, gains a ton from electricity.
You notice something different about this factory, other than the lights being a little nicer: these fellas are standing here at the end of an assembly line. To have an assembly line, I need to turn the factory sideways. We all have to be on the same factory floor. To do that, I need to locate the factory where land is really cheap. When I locate the factory where land’s really cheap, can I do that? If I can plug in every individual machine and I don’t need to be near a source of power like a water wheel — I just need electricity — then I can move the factory to a place that makes sense for this type of production.
Once I have an assembly line, I’m going to change the type of workers that I hire. Once I change the type of workers that I hire, I’m going to change the incentive system that I use inside the firm. Now we’re really cooking — now we’re getting productivity improvements. The plug-in-the-machine-that-had-the-belt-coming-down-from-the-shaft barely matters. The reorganize-the-factory-in-response-to-this-technology matters a lot.
I think this is true for essentially every new technology. So if I’m looking at AI and where it’s going to start seeing impacts on science, or productivity more generally, my thought is not, “What is the existing task content of jobs, and how much does AI affect that task content?” That’s nonsense — despite the fact that this is half the papers you’ll read on this topic. The interesting thing is: when are we going to start doing new tasks that we can only do with organizational changes in response to what the technology makes possible? How long is that going to take? What are the barriers? That’s really what we want to know, because that’s what affects growth.
Pick your favorite technology. A well-known example is ATMs.
This is the diffusion curve for ATMs. We had them in the mid-’70s. They slowly diffuse over 30 years. What happens to employment inside the firm? It goes up initially, as we move the teller slowly into higher-productivity jobs. By this point we’re totally reorganizing the type of task done inside a bank. The person who takes your money and puts it in the vault — that person’s gone by this point, the ATMs are doing that, and the people inside the bank are doing completely different tasks. But you’ll notice labor doesn’t fall, productivity still goes up. And importantly, we’re talking 30 years here.
Conditional on technical improvements, what will happen to income and growth?
This is the key question behind everything we’re going to see today. I’m going to take the technical improvements as given: AI will be able to do the things that we said before. Given that it will be able to do those things, and given that AI might also improve the method of invention — it may give me new ways to do research on robotics, for example — what’s going to happen to income, what’s going to happen to growth, and what’s going to happen to TFP?
I think this is the right model here.
This is a bad model:

This is the model like in your growth theory papers. The theory is right; it’s just — all theory is wrong, it’s for a specific use. In these endogenous and semi-endogenous models that Chad would have taught you about already, generally what happens is I’ve got some function that depends on how many scientists are working on new things, and the more scientists work on new things, the more growth I get. This might have decreasing returns to scale, but basically that’s every endogenous growth model.
I think that’s a bad way to think about AI. It’s true that we’re going to have a million Einsteins in a data center in terms of intelligence. I don’t think that’s the important factor for the outcomes we care about. The reason is that I’m going to expand the function of scientists to a more interesting model, and we’re just going to look at each of these components today.
All the existing organizations in the world have some type of structure. The management folks would call this architecture. This includes everything from fixed components that are hard to change — your brand — but also things like relational contracts inside the firm. For instance, if your firm currently finds it optimal to promote internally and use career concerns rather than bonus pay to incentivize workers, and you see AI coming and think, “We’re not going to promote internally anymore,” you’re going to find out that you can’t actually do that. Everyone’s going to quit, because they’re going to think you reneged on a relational contract. So there are many aspects of organizations that are sticky, that go beyond physical things.
Those organizations exist, however. They are not going to disappear tomorrow. They’re the ones that are going to take some existing stock of ideas and their current structure, combine them, and try to produce new ideas. Hopefully they’re going to try to modify the structure of their organization a little bit to more effectively use AI in doing this task.
Once I have new ideas, I need to diffuse them across society, or I need to extend them by creating complements. That’s what’s going to lead to growth.
You can probably already imagine quite a few pieces of this little syllogism are not necessarily affected positively by AI. Other things need to happen, or things that take time or investment need to happen, before we get improvements. That’s essentially how we’re going to structure today’s lecture.
The sticky structure of organizations
Here we’re going to build on top of Paul David. Paul David is an economic historian by trade; he does have some light models in some other papers. If I were to try to model this — how does organizational architecture limit the ability to use new tools? — I think this is the right way to model it. I haven’t seen anything that formal on this, to be honest. I’m building on a paper by Henderson and Clark in the early ‘90s that’s also not super formal. If you’re a good theorist, this is worth working on.
Because of these organizational frictions, it seems to be that with new technologies — especially general-purpose technologies — we see adoption follow essentially the same pattern in temporal order.
First, a new technology is used by one person inside an organization who can use it within their current workflow without changing anything else. You’d see something like this on AI: “I write code for a company, now I write some of my code using an AI agent.” Everyone is doing this. It requires very little change in the organization. It’s very easy to do.
If we step to the next fastest, it’s that a team within an organization adopts without changing their existing workflows. This is tougher. The key parts of how teams perform work are team incentives and costly knowledge flow. A person can adopt something without needing to communicate with anyone else, or without worrying about team incentive problems. A team doesn’t have that same luxury. If I’m writing computer code as a team, I probably want to also change how I do peer review, how I do unit tests — checking whether the code worked — and how I do evaluation and verification. In order to adopt an AI-first workflow as a team in writing code, I also need to adopt changes to those other practices. Some of those changes in practices might mean the kind of people I hire on my team will be different.
This is where you hear — with a lot of caveats — that a few research papers have found a decline in the hiring of young people inside firms. That’s being driven by this team adoption. We’re starting to change the nature of how we set up our team in response to AI. I don’t think those papers are totally right, but I think that’s the right intuition.
Third thing — this is when it’s really hard. This is the point where, with AI, basically no one’s here yet, but with old GPTs we eventually get there. The organization needs to change some architectural feature in order to adopt. Usually this means they’re doing some new task, not an old one. If you think about the kind of things an organization does, they take the existing technology, preferences, and constraints, and in equilibrium they maximize given their environment. We drop this new change in a factor price — this new change in the price of predicting things with AI — and the type of tasks I should do should be different. The problem is, when the factor price changes, as we go from the short to the long run, I also need to change the mix of capital and labor. I need to change the way I incentivize workers internally. I need to change the type of talent that I hire, and so on. This is fairly difficult. Now we’re getting to the point where I’m moving the factory with electricity out of New York to the suburbs of Detroit and turning it sideways. Now we’re getting pretty big improvements.
The option value bites — it’s costly to make these architectural changes, so you don’t want to do them 10 times. There’s some benefit, if there’s uncertainty in how the technology is going to develop, from waiting until you’re more certain about what the capabilities of the technology will be before paying this cost of reorganizing your firm.
The fourth one is really hard. Sometimes I would like to do stuff that requires regulation to change, requires partners in my supply chain to do things differently, or clients to be able to adopt a technology before I can sell them the thing. These are really difficult. My favorite example of this one is this particular fellow.
If you’re an economic historian, this is a legend from rural North Carolina. This is Malcom McLean. He invented containerized shipping. Containerized shipping is a little bit similar to AI when you think of the frictions. Imagine what it takes to get a productivity benefit from containerized shipping. It’s not just the idea that I can move the shipping container right onto the truck. I’m only getting the benefit once:
The port is set up for this,
The unions are allowing it,
I have the cranes that pick up the box and move the box onto the truck
The trucks are built to drive more efficiently given exactly the size of container; and,
The ships are built to stack containers exactly.
Once all that happens, then we get a massive productivity gain from containerized shipping. Until that happens, putting your stuff into a box of this particular size barely matters for productivity.
We’re in the era of AI, just like any new technology, where we’re a long way from the containerized-shipping world. Even worse, the capabilities of AI are changing and improving over time — the lines are going up on those graphs I showed you, which was not true of containerized shipping. What it would be able to do was identical in 1970, 1980, and 1990. That’s step one, I think, on why we’re going to see slow diffusion of AI, and slow productivity growth in science.
The productivity J curve
There’s one great paper here by Erik Brynjolfsson, Daniel Rock and Chad Syverson.
They wrote this less about AI than about information technology. They call this the “J curve.”
Measured productivity looks really bad with all new technologies. Why is that? In order to make these architectural changes, I need to invest some money today. I’ve still got a bunch of workers. Imagine the worker that was using the old technology now needs to learn how to use AI, so for the next six months they don’t produce any output, they’re just learning how to use the AI. Or I’ve got duplicated effort as I set up teams to experiment with using AI inside my firm. Now the inputs stay really high, the outputs fall — what happens to productivity? It goes down.
But it’s only because I’m measuring the intangible capital I’m developing incorrectly. What that worker is learning should count as capital owned by the firm. If I properly measure intangible capital, I’m getting the productivity shock from the start. What they do in their paper is run Tobin’s Q-type finance regressions and try to measure intangible capital. What they see is that the firms that get big benefits from digitization later on turned out to have very high levels of intangible capital very early on, which made them look unproductive because we didn’t include the intangible capital when we were doing our productivity regressions.
You might imagine something like this is going to happen also for AI, where we actually get a short-run decline in productivity because of the need to retool. How true is this? The interesting thing, for a PhD student especially, is where the open questions are.
Do other big general-purpose technologies have J curves or not? We don’t know — no one’s measured it. I suspect it’s probably a general trend, but I don’t know.
This little model I just gave you suggests that new tasks matter more for the productivity benefits of a general-purpose technology than substituting for labor and capital in existing tasks. From my read of history, if I look at the chemical industry, the industrial revolution, or digitization, I would be shocked if this wasn’t true. That said, I don’t know of any good evidence on AI, or outside of AI, that this pattern exists. Some stuff by Steve Klepper comes closest.
I think there’s a lot of nice empirical work that could be done here, looking at the extent to which new tasks really matter versus old tasks. Maybe the closest we have to this is that David Autor and friends have a really nice paper where they look at the number of new jobs that appear over time. It turns out, decade by decade, a quite large fraction of job titles you’ve never heard of before — they didn’t exist at the start of the decade. In some sense this is some evidence on what the turnover of new tasks is. But it doesn’t tell me specifically, for a given technology, how often we get new tasks.
AI in the decision problem
We’ve seen the sticky structure. What can I do, or how should I change that structure if I can make it non-sticky? Before I get there, I want to talk about what AI does in terms of creating ideas in the first place — literally, what can it do?
Another way to put that: let’s throw AI into a decision problem. We’re going way down to microeconomic theory here. There is a very simple paper by Ajay Agrawal, Josh Gans and Avi Goldfarb, related to a book they wrote called Prediction Machines. It’s one of those books written by economists that literally everyone in Silicon Valley reads at some point.
In their model, there are a couple of states of the world, and I don’t know which state is true.
I can take a safe action, but in State 2 of the world I should actually take a risky action that has a very high payoff. In State 1 of the world I shouldn’t take the risky action — it has a very low payoff. The expectation, if I don’t know what state of the world we’re in, is that I just play the safe action.
So what’s AI going to do? AI is going to tell me what the state is, with some probability. It can predict. Fundamentally, that’s what all machine learning models are doing: they’re taking in data and making a prediction. It’s a non-parametric prediction on a high-dimensional surface, but they’re making a prediction. When they make that prediction, the question is: do I want to act on that? Is it a valuable prediction? What should I do in response to it?
Ajay, Avi and Josh call this judgment, which is basically: I know my utility function. There’s knowing the state, and knowing the utility function. I take an action based on the combination of the state and the utility function. This means that across a wide part of the parameter space, AI predictions and human judgment are actually complements. As AI gets better, we’re lowering the cost of prediction. If it’s a complement to judgment, this raises the value of judgment. Knowing what to do conditional on a prediction becomes very valuable. Being able to make the prediction becomes much less valuable.
The canonical example they give of this is: take a cell phone that’s got a map on it. Up until very recently, if you wanted to be a cab driver in London, you had to study for a couple of years, know every single street, and pass this exam on something called the Knowledge. It was a very challenging exam, because London’s got all these little curvy streets — you’ll never find anything. Until you have the Knowledge, you can’t be a taxi driver. But once we have a smartphone, the Knowledge — the prediction of where to go — is: I just press a button on my phone and it tells me, even controlling for the traffic. That cost of prediction has gotten very cheap. What’s gotten very valuable is anything that’s a complement to that prediction.
We can do things within our existing system, like hit less traffic as a London cab driver — that’s mildly interesting. Or, once I get to the point where I can redesign the organization and get around organizational frictions, I realize I can just have people coming home from work pick you up in an Uber and drive a taxi for two hours with no training at all. I can redesign the entire taxi system. That’s the mapping from the decision problem into the organizational structure.
So AI is a drop in the cost of prediction. You might wonder about judgment: are humans that good at this?
We held a conference in 2017 — it was the first NBER AI conference — and Dan Kahneman was there. It was a crazy conference. I think we had 10 Nobel prize winners that attended this conference, out of about 50 people, and half of them won it since the conference. We invited the computer scientists and we invited the economists. Kahneman was giving a comment on a paper about decision-making, and he was trying to think, how should I think of human errors? Normally you think of human errors in behavioral economics as bias. But he’s like, “You should think of it like noise in the prediction function — random noise.” It’s not that the humans are necessarily in an agency problem or face some kind of behavioral constraint; it’s just that prediction is hard, and we don’t do a good job of making predictions often, as humans. He tries to take seriously this idea that it’s not that AI will have more or less bias — it’s that AI will have less noise. That’s the interesting thing about it, and he tries to draw the implications. It’s very consistent with the idea in Agrawal, Gans and Goldfarb.
If I want to map from “I have AI, it’s used in a decision problem, the decision problem is used in an organization, the organization has some frictions, given those frictions it translates those decisions into some outcome like a new idea, that outcome then diffuses, then we get growth” — I think that’s the right way to think about AI for science. We have to understand that first problem, the microeconomics of how AI feeds into a problem.
Since this paper, there have been quite a few other papers. With Susan Athey and Josh Gans, we solved a 1980s-style delegation problem: when should the AI make a decision, when should the human, and who should have the final decision rights? The interesting result there is that you may want the AI to not give its most accurate prediction to the humans all the time, because humans face agency problems.
Imagine I’ve got a truck driver down the road. 99% of the time the AI makes a correct prediction on how to drive, but I want the truck driver to pay attention in case unusual events happen. If I just give you that prediction, that’s going to dull the incentive for the truck driver to pay attention, because the probability they crash goes way down. What I want is to choose the prediction by the AI that maximizes the — conditional on the human not getting in a car crash — paying-enough-attention probability. That’s essentially what we’re going to solve for. It turns out that if I’m not just delegating but I’m designing the mechanism, I can actually do better: I don’t have to make the AI be wrong. I can just use full information, probably not surprisingly for those of you who know your micro theory.
This paper by a different Agarwal — Nikhil Agarwal, Moehring and Wolitzky — gives a really nice sufficient statistic for AI-human collaboration and shows how it applies across behavioral agency settings.
I have a paper with Josh Gans on: given that humans and AI work together on a problem, do these benchmarks that we’ve set up make any sense anyway? Surely what we want is the benchmark that maximizes the value created by the human plus the AI together, in whatever organizational setting they’re in. We show how you might be able to measure that and design better AIs conditional on their use inside organizations.
There are a lot of papers trying to get at this idea of what exactly AI is doing. If you’ve never seen these papers, your starting point should be: AI lowers the cost of prediction — and then follow from there.
AI and organizational structure
We know what AI does in the decision function. We know that organizations have sticky structure. So what should happen over time to organizational structure?
There are a couple of papers on this that I like. This paper by Ide and Talamàs is building on a Garicano-style model. That model is: there’s some continuum of possible problems. If I’m 0.8 smart, I can solve any problem between 0 and 0.8. If you’re 0.9 smart, you can solve anything between 0 and 0.9. It takes some amount of time to solve a problem, and it takes some amount of time for me to ask someone what their answer to a problem was — though less time than it takes to solve it.
With just that simple structure, I’ve generated the entire corporate form that everyone uses, with reporting structures. No agency problems at all need to be assumed here. I’m going to put the problems that everybody can solve, that are kind of trivial, at the bottom of the structure. I have all the people who can only solve problems up to 0.3 work on those. Then, as those people are working on problems and they get something that’s harder — say a 0.6 — I have them kick that up to me, their boss, and I solve that. I hire just enough people that the number of problems that get kicked up to me, and the number of explanations they tell me of the stuff they solved, exactly fill up my day.
Imagine we’ve got more and more problems and I still can’t solve them. I’m going to have someone above me who can solve from 0 to 0.8. My kid will kick up the 0.7 to me. I realize I can’t solve it, so I’m going to kick that one up to my boss and they’re going to solve that one. Now we just get this nice structure. That Garicano model tells me exactly how many employees a given manager can manage, what happens if the cost of communications falls or rises, and so on. The model explains the world with humans.
In this Ide-Talamàs paper, we’re going to drop AI into this structure.
What can the AI do? Let’s say it can predict the answer to any problem in a given set — maybe between 0 and 0.4. We’re also going to assume that all AIs are the same, so you can ask a million AIs — it doesn’t matter; they can either solve it or not. We’re going to assume the AIs are relatively cheap — you’re just paying for the API. The final thing is that any person, in principle, could query an AI, so it doesn’t have to be just the boss, it could be anyone. Sometimes we’re going to let the AI make autonomous decisions, and sometimes — for legal or other reasons — we’re going to have to have a human spend some time implementing the AI’s decision.
Let’s combine all that and see what happens. We start with this baseline Garicano model.
I know how many workers I’m going to hire — the better workers I have, the less they escalate, so the more I can manage, because it takes less time to manage them.
Now I drop the AI in.
What’s going to happen? It turns out to depend on how good the AI is, and on whether the AI can be autonomous. Let S be the set of managers. If the AI is as good as my baseline workers but not as good as the managers — it can do some of the tasks my workers can do, but none of the tasks that I can do that my workers can’t do, this is really bad for the low-level workers. I’m going to fire some of the junior workers and query the AI directly for the answer for these easy problems. My wages are going to go up; I’m super productive.
On the other hand, if the AI gets a little bit better, so now it’s middle-management quality in the type of problems it can solve, it actually starts doing the opposite. These middle managers were being paid a lot for their unique ability to do this work; the low-level people, I just needed time to throw questions down to them. So we’re going to get rid of these expensive workers now and rely on the lower-level workers to clean things up, or do the complements only humans can do if AI can’t act autonomously.
I think that’s a good example — good enough, it’s in the JPE — of someone taking organizational structure seriously, seeing what AI is going to do to who does work and who earns wages in general equilibrium, and just solving out the model.
This paper doesn’t have any innovation in it — it doesn’t have any science. So I actually think it’s pretty open, both to map this paper’s results empirically into the structure of scientific labs, for instance, and it’s quite open even just to think about the theory of what might be different. In this world the type of tasks that the organization does is fixed, but you can imagine a world in the sciences where the type of things I do look wildly different. If I have AI to do a bunch of stuff — we spend zero time inverting matrices now. The structure of our work looks very different, because inverting matrices is free. Adding those complications, I think, gets this a little bit closer to being useful.
A similar one, on how AI might change organizations, is from Shahidi et al. Their question is interesting: if you know your theory of the firm, you might be asking, why do firms exist at all? We’ve got a number of theories — residual control rights theory, resource theory, and transaction cost theory. In a transaction cost theory of the firm, what I’m doing is balancing bureaucratic costs versus the transaction cost of operating in a market. It’s somewhat difficult to gather prices, find customers, haggle, and all these things I need to do in the market, that might be cheaper inside the firm. But inside the firm it’s quite bureaucratic, and I’ve got coordination failures. So we’re balancing those two.
This is an interesting paper that says maybe the big effect of AI, at least in the short to medium run, is reducing transaction costs. If I want to buy something online, I could try to search for the lowest price, but realistically I’m not going to do it — I’m going to go to Amazon. If I could just ask my agent, “I’m buying this shirt, find me the best price that ships quickly and is reliable,” it’ll do something for 15 minutes and then buy the shirt. Now the transaction cost of doing that search is zero. If you imagine all sorts of other things — we’re already seeing AI used in monitoring of workers everywhere; simplifying bargaining, doing search, enforcing contracts — in principle, this should reduce transaction costs, allowing me to do more things outside the boundary of the organization.
This is going to matter a ton for wages — I’ll let the labor folks worry about that. But for science it also matters a ton. Right now transaction costs seem to be quite dire in limiting the ability to do large-scale science across organizational boundaries, either within the academic or the private sector. To the extent that we can verify you’re doing something useful for my firm earlier, it might open up all sorts of new scientific institutions that can contribute to innovation.
I wouldn’t say AI as a reduction in Coasean transaction cost is necessarily the most interesting thing about AI, but I do think this is a good example of taking the organization seriously, rather than just assuming I have some fixed number of scientists and that’s going to translate into growth in some way.
I would also say — combining these papers — I think this is what people are worried about in some sense, but I think it’s a good thing for science. These two results combined mean that people who have good taste as scientists — in research questions, on what to do — and are able to have a greater span of control — because they use AI for some tasks, or AI can manage other AIs — have very high marginal product.
You always hear one million Einsteins in a data center — not one million scientists. Part of it is that we think there’s a big difference between Einstein’s ability to identify interesting scientific questions, and to know it’s worth continuing down this road or not, compared to the average scientist. The average scientist does what Einstein says, and not vice versa. So you might wind up with some really interesting new organizational structures.
If you want to see some empirical papers on this that are already floating around: AlphaFold, DeepMind’s protein-folding results — AI for protein folding dominates prior approaches to understanding protein structure. We have a few nice papers:
Carolyn Stein and Ryan Hill have a nice paper that just came out.
My student Gabriel Cavalli has a nice paper.
I have a few papers on AlphaFold in particular, because the shock hit a few years ago.
I think you’ll start seeing a lot of “organization of science” papers in the near future, because the optimal organization of science in response to this AI shock — there’s no way it looks like the status quo. You might think especially sectors like academia are very sticky. There’s a paper in NBER that shows a top-percentile-by-citation AI researcher — the median salary in the US is $1.8 million, in the private sector. In the public sector it’s $360,000. Why? Because the productivity of these researchers has gone way up, but academic salaries at a department change very slowly. Now we’ve opened up a 5x gap, and we can’t retain these people. Apply that stickiness to all other parts of the academic organization of science, and you’ll start to see where there are really interesting inefficiencies beginning to arise because of organizational structure.
New ideas
What new ideas are AI going to help with? There are quite a few papers on this. One I like quite a bit is by Ajay Agrawal, John McHale and Alex Oettl. What they’re thinking about is: I have a bunch of experiments I can run. Those experiments generate information about what the next experiment I should run is. I have a limited amount of time. As a human, of these million things I can do, I have some ability to rank them, but I rank them imperfectly.
What the AI can do is tell me, given where I am in this decision problem, which are the highest-value experiments to run in the short run. AI is not going to run the experiment. We’re not having the robots carry the pipette across the lab — although, if you haven’t seen an automated lab, I would recommend visiting one, because the robots do in fact carry the pipette across the lab. But we’re just going to use the AI to rank designs.
What’s going to happen is: I’m going to propose, or the AI is going to propose, what I should work on.
There’s some exponential hazard rate that I find the thing that I’m looking for. Once I find it, or if I decide it’s not worth working on anymore, I can start working on some alternative. Just this reordering of the things I work on, without changing anything about what I could have worked on, is actually substantial, for reasonable parameters, in its effect on the number of useful new ideas I get.
I think this theoretical idea is actually all over the place, including with skepticism. If you’re an empiricist, look at something like AI for drug design or drug targeting. There’s a ton of well-funded companies working on this problem. There are a lot of people in the biology and pharmaceutical world who don’t think this really matters that much. They think they have way too many targets already, and the problem is it just takes too long to run things through trials. I guess if I could find something that was more likely to get through trials, that’d be useful. Can AI help with that? A little bit. It’s not totally clear we’re at the right stage. But at least theoretically, if we were able to not start a billion-dollar trial on something that’s going to fail — it’s a huge bang for our buck in drug development, and we can move quite quickly.
This is a nice paper, because the AI doesn’t just raise the productivity of biology researchers. The AI helps solve a specific problem that I can map into parameters of existing models, and I can try to put some empirics onto this question. That’s what we want.
Bottlenecks and the Baumol effect
I just told you that the pharma people think we have a bunch of targets. The bottleneck is getting them through trials. Running the actual trials is very expensive. If AI can reduce the cost of running trials, that’s actually more important.
Whenever you have AI improve on some part of a production function that involves complements — something that’s in a nice Cobb-Douglas form — that actually just makes the bottlenecks more important. You probably know this as a Baumol effect, after William Baumol. He said it doesn’t really matter how rich we get: to play a barbershop quartet still requires four people playing instruments. You can be as rich as you want, you still need that. Even worse, I need to get those four people to not work on the high-productivity things they could work on, because our economy’s grown, so I had to pay them more than I would have paid them back in the day, even though they’re doing the exact same work.
When I’ve got these complements, the relative importance of the remaining bottlenecks becomes increasingly important over time. It’s much better, for growth, to move through all tasks equally, rather than to move through one task to the extreme and another barely moves at all. I think this is a primary driver of economists’ skepticism toward Dario Amodei saying we’re going to have 10% growth per year. After a few years of 10% growth, I’ve now bottlenecked something else. Unless I solve that bottleneck — what are you going to do?
There’s a formal model of this that I quite like, from a paper by Ben Jones. Ben wrote a really nice paper that gives me some empirics to look at. R&D — there is a continuum of things I can do. Some the machines can do; some the humans can do. We’ll make the machines as good as you want at doing their tasks, and we’re going to see what happens to growth.
There’s the productivity of the machines, m.
The fraction of that continuous task that AI can work on is γ.
The bottleneck strength — meaning how complementary production is — is going to be θ, just like in a CES production function.
What we’re going to do is: you tell me, this is my research objective, these are the things that are necessary in order to do the task. I can estimate from current data what θ is — how complementary they are. I can estimate from your projections what AI can do — what m is going to be and what γ is going to be. Now I can just tell you what growth is going to be, at least in this sector.
In this model, the task-level output — per task, on things I’m going to use the AI for — is going to be the quality of the AI times the current state of that task. On things I’m going to use the human for, I have to use some amount of human labor in order to do that task. Let’s use a nice CES-style aggregator.
What I’m going to do is look at the total contribution of the machines to growth. It’s going to combine the machines’ quality across these tasks that it knows how to work on, and then adjust for the fact that it can’t work on some tasks. As long as the machine rent is not too high, then we’ll use the machine — but it doesn’t really matter, because I’m going to let the machine quality go to infinity.
Now I can just put numbers on this, and it’s quite surprising.
The rate of growth depends on the machines’ expenditure share in R&D, and the expenditure share turns out to be really constrained by this complementarity. I get really good at inverting matrices, but I still — when the mouse escapes that I inject a cancer into — I still have to run across the room and catch the mouse as a human. A little bit more improvement in the speed of inverting the matrix barely matters, because the whole constraint is: every time a mouse runs away, I need to catch it.
To give you some numbers: let’s imagine we have a nice harmonic average across tasks — that would be θ equal to −1 in a CES. The overall productivity is a harmonic average of my quality on each individual task. Then, with reasonable assumptions, letting infinite quality for the machines, we wind up getting growth of 2.25 times today’s productivity. So, 10 times intelligence, 100 times machine intelligence — it’s not 10 times growth, it’s not 5 times growth, it’s 2.25. The only real way out of this is that we kill off all the bottlenecks. Even a small number of bottlenecks on machine intelligence in the science production function starts to bite really hard.
Conclusions
I hope you can see there’s just a ton of open questions.
I’ve got these existing organizations. How quickly are they going to be put out of business by more productive firms? Normal competition policy, IO.
When the new firms do things, how are they going to use AI? That’s for you micro theorists, in the decision function.
How am I going to incentivize people, and how is that going to map up into the structure of the new organization? That’s for you organizational-econ, IO folks.
Given that new organization, and its mixture of humans plus AI, and taking the projection seriously about what the future looks like from the computer scientists — how is that going to map into new ideas in each sector? How is that going to map into growth?
What can we say at each stage of those things empirically right now that’s informative, given those theoretical models of where we’re going to wind up in the future?
If I were a student right now interested in innovation — hopefully I’ve convinced you that a lot of the interesting innovation questions are going to touch AI. Those lines on the graph we started with, they’re still going up. The million Einsteins in the data center will happen. It’s just that the implication of that might not be what a lot of folks who aren’t economists think, and we have a ton of scope for helping people understand things — even understand that in the short run we might see declining measured productivity because of that J curve I showed you. It’s quite an interesting time for this type of research.



























