When I started my own business, I was back in a zero-to-one build. It wasn’t my first. I had built projects and functions from scratch throughout my career. Yet, that experience helped less than you’d think. Mostly, knowing what comes next created impatience to skip ahead to it.
The company hadn’t found product-market fit yet, and my stint in tech had ingrained an appetite for efficiency and scale. My urge to build the machine, even before I knew what the machine was supposed to make, was strong. But you cannot standardize something that hasn’t found its shape. This piece is about that tension: systems need variability while they’re finding their shape and consistency once they have one. Who, if anyone, decides how much variability a system keeps, and when?
It’s a tradeoff I’d been on the edges of throughout my career. At the fifty-year-old consulting firm where I started, the variability was the product: rigorous, tried and true methods, but no two engagements applied them the same way. At the government agency, consistency came first, because an inconsistent decision, in benefits, in enforcement, in due process, costs more than a slow one, and some of that inefficiency was doing constitutional work, protecting individual rights and preventing the consolidation of power. The nonprofit in between ran like a business, but with no market to check its setting. Each of them had a tolerance for variability, and none of them chose it deliberately. It was habit, survival, or instinct.
Then tech, where I joined one company in late Series A and stayed through late Series E, and watched practices that had been the key to success in one stage become the Achilles heel of the next. Everything about each of those places had been fit for purpose, at least at some point in time. Some of the standardization calls were clearly right. Some were made too early, before anyone had learned enough to know what should be scaled.
To name my bias: I love high-variability contexts. Building something from nothing, testing and experimenting, sorting through many wrong answers before finding a right one. So I start from the position that variability is the fuel adaptive systems need to innovate and create, not noise to be standardized away. On the other hand, too much variability creates risk: inconsistent quality, inconsistent decisions, inconsistent experiences.
Tolerance for variability is a setting, not a virtue. Genomes have one. Societies have one. Companies have one. And every one of those settings was calibrated by an environment that may no longer be the environment in front of it. Having the wrong setting from time to time is inevitable; environments change faster than calibrations do. Where we go wrong is not knowing we have a setting at all, and therefore never checking whether it still fits. I won’t arrive at a settled answer about where anyone’s setting should be. What I’m after is a better way of asking.
What biology knows about this
Before I changed my undergraduate major to psychology, I studied evolutionary biology, and the question that pulled me in, nature and nurture, is largely a study of variability: where it comes from, what it does, and what we lose when we optimize it away. Natural selection needs variation to work on. A population with no genetic variation has nothing for selection to filter, and evolution stalls until new variation shows up. Biology treats variance as a strategy. In environments that swing unpredictably, a genotype that produces one “best” phenotype is betting that the future looks like the recent past; a genotype that produces a spread of phenotypes does worse on average right now and survives regime change. Desert annuals hold a fraction of their seeds dormant even in a good season, insurance against the drought that hasn’t come yet. Bacteria facing antibiotics keep a small subpopulation of slow-growing persisters that survive the assault and reseed the colony. The organisms that got it wrong don’t exist to study.
The alternative strategy has a name too. In 1942, Conrad Waddington called it canalization: development buffered so well that it produces the same outcome no matter the noise. Buffering is useful, which is why it evolves. The catch is that a trait that stops responding to disturbance also gives selection nothing to work with when the environment shifts. And suppressed variation isn’t always gone. Sometimes it’s destroyed. Sometimes it’s stored where it can be recovered. That difference will matter again when we get to AI.
In evolution, the variance often lives at the population level, and a skeptic could say the market plays that role for companies: over-standardized firms dying and being replaced is the system working as designed. But the people inside a firm cannot diversify themselves across ten thousand startups, and the market’s bet-hedging is cold comfort to the company that turns out to be the losing bet. If you run a firm, you are not the population; you are one of the seeds, and you get one draw.
A global stress test
Between genomes and companies sits a level where this trade has been measured across dozens of countries and then stress-tested by a global event.
In undergrad, I took cross-cultural psychology with Dr. Michele Gelfand. She studies how strongly societies enforce their norms, what she calls cultural tightness and looseness, and her research across dozens of nations shows the setting isn’t arbitrary; it tracks threat history. Societies that have faced more disasters, disease, scarcity, and invasion tend to have tightened, because coordinated compliance is what gets a population through an acute collective threat. Societies with gentler histories loosened. Tight cultures get order, coordination, and self-control. Loose cultures get openness, tolerance, and creativity. She is emphatic that neither is better, and the pandemic gave the theory an unusually direct test.
In 2021, in The Lancet Planetary Health, Gelfand and colleagues examined 57 countries and estimated that loose nations had roughly five times the COVID cases and nearly nine times the deaths of tight nations by October 2020, controlling for wealth, inequality, density, median age, government efficiency, underreporting, and more. Brazil and the United States each passed 24,000 cases and roughly 700 deaths per million in that window. Taiwan had 22 cases per million. The evidence is correlational, and tight countries tend to have other advantages, state capacity and recent epidemic memory among them. But for a threat that demanded fast, uniform, population-wide behavioral change, the pattern is what the theory predicted, and the margin was large.
If the story ended there, it would be a case for tightness. It doesn’t end there.
A small analysis in Psychological Medicine that same year found that across twelve countries, cultural tightness correlated negatively with willingness to receive the COVID vaccine, and willingness tracked how much virus was actually circulating. Tight cultures had contained the virus so effectively that their people perceived little risk, and vaccine demand fell. The setting that won the containment phase was mismatched to the vaccination phase, within the same crisis, within the same year. Later work found the mirror image: as conditions eased, looser cultures relaxed their guard fastest, while tight cultures held their vigilance well past the point the data justified it.
The pandemic’s lesson is more unsettling than “tight beats loose.” The same looseness that cost the United States dearly in 2020 is tied to the openness that produces its disproportionate creative and scientific range, including, not incidentally, several of the vaccines that ended the acute phase. Nobody gets to be maximally both. Each setting was right for some environment and will be wrong for another, and the settings persist long after the environments that calibrated them are gone.
Organizations have tightness too. It shows up in how strictly process is enforced, in what actually happens to the person who does it a nonstandard way, in how much consequence attaches to deviation. A company that tightened during a compliance scare, or a scaling push, or a near-death experience, and never revisited the setting, is running a calibration built for a threat it may no longer face. I have worked inside that company. So, I suspect, have you.
The newest lever on an old dial
Organizations have been adjusting this dial deliberately for over a century. Frederick Taylor’s scientific management compressed variability in physical process, replacing each worker’s judgment with one best method, and it delivered the productivity it promised. Six Sigma compressed it in outcomes, 3.4 defects per million opportunities, and Motorola credited it with billions in savings. The actual discipline was more careful than its caricature suggests; the target was unwanted variation, not variation itself, and in some domains – pay and promotion decisions chief among them – driving variance down is the right choice, because unmanaged variability there is how inconsistency and bias enter high-stakes decisions about people’s lives.
The costs took longer to surface. A twenty-year study of the photography and paint industries found that the more intensively firms adopted process management, the more their innovation shifted toward refining what they already knew, at the expense of exploration, the pattern James March named exploitation crowding out exploration. 3M ran the experiment on itself: a new CEO put Six Sigma through everything, including the research labs. Efficiency improved, the pipeline of genuinely new products thinned, and his successor pulled the program out of R&D, explaining that “invention is by its very nature a disorderly process.” He wasn’t rejecting Six Sigma, which stayed in manufacturing where it belonged. He was drawing a boundary around it. Every tool in this lineage worked where work had found its shape, and cost something where it hadn’t.
AI is the newest lever on the dial. It is the fastest and cheapest one ever built, and it works on a different target. Taylor compressed physical process. Six Sigma compressed outcomes. AI compresses ideas.
The evidence so far is early, and mostly pointing the same direction. A 2024 experiment in Science Advances found that writers given AI-generated ideas produced individually better, more creative stories, while the stories across writers became more alike. Individual writers gained while the group converged. A 2025 study of roughly 2,200 college admissions essays found that each additional AI-written essay contributed fewer new ideas to the pool than each additional human one, even after the researchers tuned the prompts specifically to force diversity. The mechanism is simple. Models tuned on aggregated human preferences learn to produce what most people rate as good, which pulls output toward the center of the distribution. The center is a fine place for a first draft. It is a poor place to find anything that could differentiate you.
This isn’t an anti-AI argument. I use these tools daily. But I think a lot about how I use them and whether it’s bringing me to the center or pushing me to the edge.
First, unlike canalization in biology, this compression is an engineering artifact, not a law of nature. Output diversity responds to how the tools are built and how they’re used; in one large experiment, people who encountered AI ideas as stimuli to react to produced a more diverse pool of ideas, not a narrower one. The research on training models on their own outputs, the model collapse result, shows distributions narrowing from the tails inward, but follow-up work shows the effect is largely avoidable when fresh real data keeps flowing in. That is the banked variation from the canalization story: collapse comes only when the reserve stops being replenished. These are choices someone is making, at the model labs and at every desk where the tools get used, and choices can be revisited.
Second, the gains are measurable, and they are table stakes. In a study published in the Quarterly Journal of Economics in 2025, customer service agents with an AI assistant resolved about 15 percent more issues per hour, and the least experienced gained closer to 30. Every competitor can buy the same lift. Dave Ulrich puts the logic in financial terms: performance is revenue over costs, and AI-generated capacity works on the denominator. That work is worth the effort, but costs have caps while revenue has no limits, and cost efficiencies tend toward parity as competitors match them with similar technology. Revenue comes from differentiation, and differentiation runs on the kind of output these tools compress by default. A company that adopts AI and stops at capacity hasn’t gained an edge. It has paid to arrive where everyone else is arriving, slightly faster.
One more thread worries me most as a psychologist, because it compounds. Judgment is built through repeated exposure to ambiguous cases. The novice gains in the productivity studies come from AI substituting for judgment the novice doesn’t have yet. If AI increasingly intercepts the ambiguous cases before people have to sit with them, we’re cutting off the experiences that build judgment at the same time demand for it is rising. I don’t know when that bill comes due. It’s more likely to be years than quarters. But it lands on the people who were supposed to be ready for the decisions no model has seen, and never got the practice. Adoption dashboards won’t catch it.
The instrument, not the answer
I told you I wouldn’t arrive at a settled answer, and I haven’t. What I have, after running this question through evolution, cultures, factories, and models, is a better instrument for asking it. It has three parts.
Is variability treated as a feature or a bug? That’s our setting, and the honest answer isn’t the stated one. It’s what actually happens to the person who does it a nonstandard way.
What calibrated it? A funding crunch, a compliance scare, a founder’s last company, fifty years of bespoke work. Name the environment your setting was built for, and you can ask whether it’s the environment you’re in.
Which variability are we giving up right now, and is it the kind we mean to give up? Variability in how a routine task gets executed is usually safe to compress, and AI’s gains there look close to pure upside. Variability in judgment applied to people, compress that too, on purpose, for their protection. But variability in ideas, in the range of framings a group can generate, in the ambiguous cases that build the next generation’s judgment: that is the reserve everything else eventually draws on, and it is what the newest tools compress.
This sorting is also the beginning of an AI policy. The first bucket is where AI-assisted work should be default-on. The third is where it should stay default-off, or at least deliberately staffed with people who still need to build the judgment. Clay, a software company, wrote an AI writing policy on exactly this instinct: AI is welcome for brainstorming, drafting, and proofreading, but “writing is thinking,” every sentence has to remain the author’s own, and outsourcing the writing to skip the thinking defeats the point.
I’ve seen what this looks like first hand. During one company’s hypergrowth phase, we were pushing acquisition costs down by standardizing and scaling parts of the customer experience, and one instinct kept resurfacing: take the white-glove experience we’d built for executives and extend it to every customer. Sometimes that was the right call. Sometimes it was a timing problem. It was just too soon; a bespoke experience that fit perfectly in one context hadn’t been battle-tested for broad use. And sometimes the scarcity was the point: opened to everyone, it stopped feeling special and started feeling commoditized. Each time, the difference was what the variability was doing for us, and often, no one was asking.
Every organization I’ve been part of made these choices. Most of them weren’t making them consciously. Biology has been running this experiment for a few hundred million years, and its track record with systems that optimize their own variation away is not encouraging: they get efficient, and then they get brittle, in that order.
Organizations swing on this dial: tighten until something breaks, loosen until something else does. We usually narrate each swing as a correction, a strength overused, a miscalibration finally caught. Maybe. Or maybe it’s organizational physics: the tide goes out because it came in, and systems get thrown from exploitation back toward exploration only when the context demands it. I’d like to believe in a third possibility, a fit-for-purpose balance that can be designed and recalibrated as the environment moves. I’ve watched the swings from inside five organizations, and I can’t yet tell you which of the three is true. Every version of the answer starts in the same place, though: knowing you have a setting at all.
References
Ashkinaze, J., Mendelsohn, J., Qiwei, L., Budak, C., & Gilbert, E. (2025). How AI ideas affect the creativity, diversity, and evolution of human ideas: Evidence from a large, dynamic experiment. In Proceedings of the ACM Collective Intelligence Conference. https://arxiv.org/abs/2401.13481
Balaban, N. Q., Merrin, J., Chait, R., Kowalik, L., & Leibler, S. (2004). Bacterial persistence as a phenotypic switch. Science, 305(5690), 1622–1625. https://doi.org/10.1126/science.1099390
Benner, M. J., & Tushman, M. (2002). Process management and technological innovation: A longitudinal study of the photography and paint industries. Administrative Science Quarterly, 47(4), 676–706. https://doi.org/10.2307/3094913
Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at work. The Quarterly Journal of Economics, 140(2), 889–942.
Doshi, A. R., & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28), Article eadn5290. https://doi.org/10.1126/sciadv.adn5290
Gelfand, M. J., Jackson, J. C., Pan, X., Nau, D., Pieper, D., Denison, E., Dagher, M., Van Lange, P. A. M., Chiu, C.-Y., & Wang, M. (2021). The relationship between cultural tightness–looseness and COVID-19 cases and deaths: A global analysis. The Lancet Planetary Health, 5(3), e135–e144. https://doi.org/10.1016/S2542-5196(20)30301-6
Gelfand, M. J., Raver, J. L., Nishii, L., Leslie, L. M., Lun, J., Lim, B. C., … Yamaguchi, S. (2011). Differences between tight and loose cultures: A 33-nation study. Science, 332(6033), 1100–1104. https://doi.org/10.1126/science.1197754
Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., … Koyejo, S. (2024). Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. arXiv. https://arxiv.org/abs/2404.01413
Hindo, B. (2007, June 11). At 3M, a struggle between efficiency and creativity. BusinessWeek.
Kirk, R., Mediratta, I., Nalmpantis, C., Luketina, J., Hambro, E., Grefenstette, E., & Raileanu, R. (2024). Understanding the effects of RLHF on LLM generalisation and diversity. In International Conference on Learning Representations. https://arxiv.org/abs/2310.06452
Ma, M. Z., Chen, S. X., & Wang, X. (2024). Looking beyond vaccines: Cultural tightness–looseness moderates the relationship between immunization coverage and disease prevention vigilance. Applied Psychology: Health and Well-Being, 16(3), 1046–1072. https://doi.org/10.1111/aphw.12519
March, J. G. (1991). Exploration and exploitation in organizational learning. Organization Science, 2(1), 71–87. https://doi.org/10.1287/orsc.2.1.71
Moon, K., Green, A., & Kushlev, K. (2025). Homogenizing effect of large language models (LLMs) on creative diversity: An empirical comparison of human and ChatGPT writing. Computers in Human Behavior: Artificial Humans, 6. https://www.sciencedirect.com/science/article/pii/S294988212500091X
Ng, J.-H., & Tan, E.-K. (2021). COVID-19 vaccination and cultural tightness. Psychological Medicine. Advance online publication. https://doi.org/10.1017/S0033291721001823
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759. https://doi.org/10.1038/s41586-024-07566-y
Taylor, F. W. (1911). The principles of scientific management. Harper & Brothers.
Ulrich, D. (2026, June 2). AI * HI for stakeholder HR: Turning capacity into capability [Article]. LinkedIn. https://www.linkedin.com/pulse/ai-hi-stakeholder-hr-turning-capacity-capability-dave-ulrich-pupgc/
Venable, D. L. (2007). Bet hedging in a guild of desert annuals. Ecology, 88(5), 1086–1090. https://doi.org/10.1890/06-1495
Waddington, C. H. (1942). Canalization of development and the inheritance of acquired characters. Nature, 150(3811), 563–565. https://doi.org/10.1038/150563a0




Outstanding analysis and insights Shonna!