How I Would Improve the Multimodal Study at KUKA
Working on the multimodal human-robot interaction study at KUKA was one of my most applied research experiences. We learned a lot about how operators respond to different feedback modalities during cobot commissioning. But looking back with fresh eyes, there are several things I'd do differently to make the findings sharper and more actionable.
1. Add a verbal confidence check after each task
After participants completed a task with a given intervention, I'd ask them to rate their confidence in using it on a scale of 1 to 10. A simple verbal check, right in the moment. This gives you a lightweight subjective measure that pairs with the behavioral data. SUS and UEQ tell you about overall experience, but they don't capture how confident someone felt at a specific point in the workflow. Confidence matters a lot in high-stakes industrial settings where hesitation can slow everything down.
2. Track error rates per task
We measured task completion and time-on-task, but I'd add explicit error rate tracking: wrong inputs, missteps, retries, and recoveries. Error data tells you not just whether someone finished, but how cleanly they got there. It also helps distinguish between interventions that make tasks easier versus those that just make them feel easier. Sometimes people complete a task confidently but make more mistakes along the way.
3. Establish a SUS baseline without interventions
One thing I'd change in the study design is running a baseline condition with no multimodal feedback at all. We compared different intervention types against each other, but without a clean no-intervention baseline, it's harder to say definitively that the feedback helped rather than just being less bad than other options. A baseline SUS score gives you a reference point to show real improvement.
4. Include a think-aloud protocol for at least a subset of participants
Adding a concurrent think-aloud to some sessions would surface the reasoning behind hesitations and errors. You see someone pause in the observational data, but you don't always know why. Were they confused by the feedback? Did they not notice it? Were they second-guessing themselves? Think-aloud gives you the qualitative layer to interpret the behavioral patterns.
5. Measure learnability across repeated sessions
We tested operators in single sessions, but in reality they'd use the system repeatedly. I'd add a follow-up session (even a short one) to see whether the interventions helped people improve faster over time. Does multimodal feedback accelerate the learning curve, or does it become noise after the first few uses? That distinction matters for long-term product decisions.
6. Add a short interview on expectations vs. reality
Before the study, ask participants what they expect from the setup process. After, ask what surprised them. This expectation-reality gap often reveals design opportunities that purely quantitative measures miss. If operators expected the robot to give clearer signals and it didn't, that's a design insight worth capturing.
None of these changes would have been hard to implement. They're mostly about adding structure around data we were already close to capturing. That's often how research improves: not by redesigning everything, but by being more intentional about what you measure and when.