4 Comments
User's avatar
Al Dea's avatar

Great piece and thank you for being so precise in identifying what I think is a huge gap in the discourse around task and the broader conversation of "AI coverage." Every time I see one of those reports I wonder to myself about the efficacy of the baseline they are using to measure and I think you nailed it with why its worth a much deeper look.

To the credit of some of those reports, I remember listening to a podcast interview with the team that does these at OpenAI and I do believe the author noted that in addition to Onet they actually do use vetted experts across those disciplines in the report to confirm the validity of the tasks, but I have a hard time believing that even with that judgement that everything across each industry maps so cleanly. But you point to a much more important question worth asking regardless of what the percentages net out to.

Also, thank you for incorporating Suzman's work into your piece. I remember reading his book back in 2021-2022 but it was a good reminder for me to go revisit it. What struck me at the time of his reflections was just how different the people in those tribe associated to their job/work. I remember the point he made about how if you were a hunter and maybe got too much meat people in the tribe would laugh and make fun of you (in a light hearted way) for being too largesse. It felt like a lot of their defaults of the tribe were set to the collective, and having a good idea of what "enough" is, and I think today, that is something that is entirely different at least in Western Society. I think its hard to envision us going back to what was, but I wish we could be more imaginative around how some of these principles for the tribe could be incorporated in a more modern way, which would in turn, reshape I think how we collectively associate to work.

Shonna Waters's avatar

Thank you for this, Al! You're right that the OpenAI team uses vetted experts to validate their task ratings, and I don't doubt that they're doing due diligence. The caveat I'd offer is that expert judgment can make the ratings better, but it can't fix what's being rated. If the instrument is a list of task statements, even perfect ratings on the target are still deficient. The judgment isn't the constraint, it's the model underneath it. That's what I found so useful about the HumRRO critique.

And I love that you brought up the meat-sharing bit! A successful hunter gets teased, the meat gets shared, which means the whole system is operating to define the values that determine status. Those values were built into daily practice through social norms. Which makes me more hopeful than "we can't go back" might suggest. Defaults are design choices. We redesign them in organizations all the time; we've just rarely asked what a modern default for "enough" would look like.

I wonder what that could mean for how we price and reward work. What would you redesign first?

Dan Riley's avatar

As always, such a thoughtful and insightful piece, @Shonna Waters. As you so perfectly frame - “Everyone's asking what AI will do to work. But almost nobody stops to define work first.” Yes, yes - a million times to that! 👏

Liz Pavese, Ph.D.'s avatar

Such a great piece, Shonna!