17 Comments
User's avatar
Luis Guirola's avatar

I haven’t read the whole column yet, but let me say that I think this effort is extremely useful. Let me tell you my story.

Earlier in my career I tried to exploit my knowledge of econometrics data science and statistics to do applied work. My statistical education went a bit beyond the standard econometrics core —I took classes in predictive modeling and Bayesian statistics with the stat folks. In the stats departament the focus was really applied, so my hope was that I would likely be able to apply those techniques to new problems.

Now, when I came back to the econ world I found that The applied people were hostile to any innovation that was not implemented smoothly in Stata (“why do you partially pool your fixed effects in a hierarchical model? This is very unconventional” “what’s the point of this Gaussian process? Just do OLS”). Nobody wants to ready the methodological section of your paper.

My hope was that I would be more successful if I shopped among the econometrics people. After all, econometrics is supposed to be about developing methods that are ready to use for economic questions, so probably my applied peers would be more receptive to these. However, attending econometric sessions in conferences was much worse: there was exactly zero focus on explaining why this was useful to solve this or that particular problem, it was all about relaxing this or that particular assumption. It seemed written for other econometricians, never with the users of those techniques in mind. So, the hope to use one these new estimators to do produce a new paper seems hopeless.

A related problem is implementation: even if they show that their estimator are useful, econometrics people rarely any software that can be used by applied researchers. This is not surprising as this is a lot of work and hard to maintain, probably not sufficiently recognized within the Econ world. Consider the big contrast with the work that Gelman and his coauthors have done with the Stan environment for Bayesian statistics.

In that respect, I think the work of new DID literature should be the standard: taking a problem in applied research, showing that the standard method delivers bad results, and underscoring how your new estimator does better; they also have developed great if R packages and tutorials and have a webpage for it.

With that backup, it is much easier to people like me to use such estimators in applied research

Paul Goldsmith-Pinkham's avatar

Yes I have grappled with my own versions of this, and agree with a lot of your sentiment. I wonder how much in the age of AI the costs of implementation will fall?

And I do think there is more focus now on the Econ side of econometrics vs the mathematical statistics

Seth's avatar

It's honestly weird how rare Gelman-type researchers are, considering how wildly successful and influential he has been as a "theorist of applied statistics". I think part of the problem is that the incentive structure of academia makes this hard to pull off as a career.

Seth's avatar

The disconnect between technical work and empirical work is, I think, pretty universal across fields--it is certainly true in neuroscience! Part of the issue is incentives and specialization: mathy researchers are people who have specialized in math, and they are evaluated by people who have specialized in math on the basis of how impressive and interesting and novel their math is. But for empirical work, you generally *don't want* super impressive or novel math; you want something relatively simple and robust and well-understood, that is *appropriate for your empirical application*.

There's a related problem, which is 'when you have a hammer, everything is a nail'. Theorists and mathematicians are strongly inclined to view all empirical problems as nails, to be whacked with whichever methodological hammer happens to be their own hobby-horse. If you really needed a screwdriver, well, you should have talked to someone in the screwdriver department!*

LLMs might be helpful in solving this problem. They have read the entire econometrics literature, but unlike a human econometrician they have no reason to push one method over another, or a more complicated method over a less complicated method.

*My own favorite thing to do, back in the academy, was talk my way out of authorship credits by convincing experimentalists not to run the complicated model they thought they wanted, and instead just run this line of R code I scribbled on a napkin. It was win-win because they got a better, clearer paper, and I got to do less work.

Per Pettersson Lidbom's avatar

A factorial difference-in-differences set-up is also discussed in the shift- share literature. For example, Kolesar discusses this set-up on p.22 in his lecture notes. (https://github.com/kolesarm/539b/blob/master/2026s_06_ssiv_slides.pdf). In his notes, he mentions some papers such as Jaeger, Joyce and Kaestner (2020) and Christian and Barnett (2024) that investigates this issue.

The factorial difference-in-differences set-up is also discussed by Kirill Borusyak, Peter Hull, and Xavier Jaravel (2025) in their paper “A Practical Guide to Shift-Share Instruments” (https://www.aeaweb.org/articles?id=10.1257/jep.20231370). See section A1 in their Online Appendix for: A Practical Guide to Shift-Share Instruments.

Per Pettersson Lidbom's avatar

I forgot to mention that it is also important to stress that under a more general fixed effect design which allows for violation of parallel trends, such as an interactive fixed effect approach as discussed by Bai (2009) and others, a factorial difference-in-differences set-up is highly problematic since it will be almost impossible to disentangle the treatment effect itself from the effects of the interactive fixed effects. Also, a factorial difference-in-differences design must deal with the problem of cross-sectional dependence since the identification is based on an event affects all units at the same time. See Andrews (2005) for a discussion of the identification in cross-section regression with common shocks.

Cameron Ellis's avatar

Please keep doing this on a regular basis!

Ben Boehlert's avatar

This seems great! One problem I often have as an early-career researcher is that I can't tell what work is important versus what isn't, and an experienced applied econometrician's take should be useful!

Per Stromberg's avatar

Wonderful - thanks Paul for doing this!

FinallyInCrypto's avatar

Thank you for this new Substack. This is a great idea. And best wishes!

Graeme Walsh's avatar

As a Central Bank economist, I find this kind of post really useful and I will recommend it to my colleagues. Thanks Paul. 👏

Manoel Galdino's avatar

This is so great. I owe you many thanks, for many different things. I watched almost all of your methods course on YouTube to study and prepare for my own course in polisci. I learned with you on using agents, reading or watching your Markus academy lectures. Your blog, your papers, now this. I’m a fan of your work.

Rebeca Huerga's avatar

Saving for later! This is gold! Thanks Paul! 👏👏👏

Will Miller's avatar

Fantastic stuff here, Paul!

Anna's avatar

As an economics master's student this is very useful ! Thank you

Cape Fear Advisors's avatar

Please keep going with this. The translation problem shows up outside economics too, and the version out here is worse, because the technical work is not distant so much as unread.

One thing worth naming about your five picks. A breakdown point for when the sign flips. Coverage that survives misspecification. An assumption made explicit that was being carried silently. Two approaches that remove a premise which was never plausible to begin with. Every one of them is an instrument for finding out you were wrong, and you selected them without saying that was the criterion. If the column has a through line, that is it, and it is a better invitation to applied people than technical merit.

On footnote 7, yes, and the precision dependence problem may be sharper there than in the teacher case. A manager's alpha is estimated more precisely the longer the track record, and track record length is not independent of the thing you are trying to measure, since it partly reflects having survived. The dependence is structural rather than incidental.