Designing the oversight hotzenplotz: as we become humans in the loop, what becomes of our work?

Designing the Oversight Hotzenplotz: What Should AI Leave for Us Humans to Do?

In our latest reading group, we discussed Klapperich and Hassenzahl’s wonderfully strange paper “Hotzenplotz – Reconciling Automation with Experience”, along with a follow-up-ish study on automated driving “Driving Hotzenplotz: A Hybrid Interface for Vehicle Control Aiming to Maximize Pleasure in Highway Driving”. For context to the readers, the original Hotzenplotz is an electric coffee grinder with a functionally unnecessary manual crank added to “create meaning”. The paper asks a deceptively simple question: when automation makes something easier, faster, or more efficient, what human experience might it remove? The authors describe this as a kind of “experiential amputation”: automation can save labour, but it may also remove agency, skill, involvement, judgement, meaning, ritual or pleasure (e.g., manually grinding coffee beans).

As a group, we found the paper playful, slightly odd, and methodologically imperfect. The language is sometimes silly, the study is small, the statistics are not very detailed, and the work clearly comes from a pre-modern AI time. But we used it as a provocation to discuss a question that is becoming increasingly urgent in human-centred AI: when AI automates (part of) a decision task, what kind of human work should be preserved, and why?

From catching errors to preserving meaningful work

A lot of work on human oversight of AI focuses on whether humans can detect and correct AI errors. That is certainly important. But our discussion kept returning to the design of work that is human oversight of AI. How do we deliberately plan and design oversight work such that while it supports the human in detecting and correcting errors, it also preserves meaning and willingness to do the work? What meaningful parts of human work might be amputated in oversight roles, and is there a Hotzenplotz design move here?

We used a worksheet activity to think through this. Lab members picked a domain of their choosing, which included radiology, peer review, research work, coding agents, and content moderation. For each domain, we went around the table and discussed aspects such as what does the AI automate, what human contribution might be “amputated”, and what would a meaningful oversight design preserve. We arrived on several interesting bits of discussion.

First, automation often removes not just effort, but also expert judgement. In radiology, for e.g., AI might automate detection, measurement, or even report generation. But the meaningful human contribution may lie in deciding what is worth highlighting, interpreting findings in a clinical context, and writing reports that are useful to other clinicians. This goes beyond information processing to finding pleasure and taking pride in exercising professional judgement.

Second, automation can threaten learning and skill maintenance. This came up strongly in our discussions of coding agents and research assistants. AI tools can make people more ambitious because they lower the barrier to trying things. Think of how many months it used to take to build a web-based research prototype five years ago, from an idea to a deployment-ready prototype, and then think about how people can replicate that within weeks today, if not days! No doubt AI tools are making us efficient, but as a byproduct, they are also reducing our need to struggle through hard parts of the work, and that struggle is often where learning happens. One of the lab members framed the ideal AI research assistant not as a tool that produces the answer but as one whose job is to improve the human researcher’s skills and abilities.

Third, we discussed how meaningful work is often social. Peer review was a good example. An AI could summarise a paper or collate reviews, but reviewing is also a form of community participation. It involves implicit judgements, venue/journal norms that often cannot be/haven’t been formalised, and most importantly, a sense of service to the research community. If review becomes fully automated, we might lose part of the scholarly community that the process helps maintain. The group also raised a similar point for content moderation. Some cases with clearly harmful content may be appropriate to automate. But many moderation decisions are about (implicit, experience-driven) community norms: what kind of space is this, what counts as acceptable here, etc. In such cases, moderation is not just an act of removing content but part of community governance and maintenance. Fully automating it likely takes something of value away from the human moderators of these communities.

A decorative Hotzenplotz crank?

A recurring concern was that “meaningful human involvement” can easily become fake agency. The Hotzenplotz crank is quite charming in a coffee grinder. But during AI oversight, an analogous decorative crank could introduce new risks. The human might be asked to approve, sign off, or “stay in the loop” without having the time, information, authority, or organisational support needed to make a judgement.

Work, labour, and the right level of automation

One of the most useful distinctions that emerged was between work and labour. Labour is the repetitive, tedious, annoying, low-value part of a task. Work is the part that carries judgement, craft, care, etc.

Of course, the distinction is not fixed. Grinding coffee may be a meaningful ritual for one person and an annoying laborious task for another. Writing a peer review may be community service for one person and a burden for another. The same task can have a different meaning depending on context. We argued that this context-dependence/context-sensitive design is crucial for AI oversight. Several examples pointed towards a middle ground. In peer review, AI might support note-taking, collation, or summarisation, while preserving reading, judgement, and final evaluation for the human. In radiology, AI might support writing clinical reports, while preserving interpretation, judgement, and diagnosis for the human. In content moderation, AI might remove clear cut harmful content, while humans can deliberate over more community-specific.

Future directions for human-centred AI oversight

The discussion left us with several open questions for future research.

First, we agreed that we need to better understand the psychological needs of oversight. The automated driving paper asks which psychological needs need to be fulfilled for driving to remain pleasurable. For AI oversight, the equivalent question might be: which psychological needs need to be fulfilled for oversight to remain meaningful, sustainable, and effective? Autonomy, competence, professional pride, learning, social connection, responsibility, and purpose all seemed important in our discussion.

Second, we need to think beyond interfaces and towards work design. Meaningful oversight requires rethinking the very nature of human-AI collaboration itself. We also need to focus on preserving the skills and motivation of the overseer. If AI helps people produce faster work while gradually eroding their ability to perform the task, then the system may be efficient in the short term but harmful in the long term.

The Hotzenplotz papers do not offer a direct model for AI oversight, but they do help sharpen the questions we should be asking. Human oversight is not simply a matter of placing a person after an AI system to approve, correct, or contest its outputs. It changes the nature of the human role: from doing the original task to judging the work of an automated system, from producing decisions to evaluating them. That shift has consequences for expertise, agency, accountability, and the very meaning of work. For human-centred AI, the challenge is therefore not only to design better oversight interfaces, but to design better oversight work: work that gives people the information and intervention mechanisms needed to exercise judgement meaningfully.

Watch this space for how we go about tackling this grand challenge.




Enjoy Reading This Article?

Here are some more articles you might like to read next:

  • Geoguessr: A new scientific frontier
  • ARC Centre of Excellence success in ASCI
  • ASCA success at ASCI!
  • Evidence-based scientific thinking and decision-making in everyday life | Dawson et al | 2024