Voice-first XR accessibility / Research through design
SpeakEasy
Designed and built a Quest 3 voice-first accessibility prototype in Unity, using Meta Voice SDK, multimodal feedback, and three research-through-design iterations to reduce dependence on handheld XR controllers.
- Operating proof
- Direct research and implementation evidence for a voice-first mixed-reality accessibility system, spanning participatory inquiry, Unity and Quest development, Meta Voice SDK integration, measured prototype iteration, and public thesis defense.
- Engagement
- Master of Design in Experience Design, SJSU
- Evidence set
- 7 reviewed artifacts
- Capability coverage
- 4 documented areas
Primary artifact
01 / 07Thesis defense
Michael Chaves publicly defended SpeakEasy as a voice-driven AI system for inclusive XR at San Jose State University on May 14, 2025.
01 / Context
Situation
•Controllers Created an Entry Barrier
The thesis focuses on people with limited grip strength, upper-limb mobility challenges, and low muscle tone who can be excluded when XR requires two handheld controllers and precise physical input.
•Voice Needed to Provide Agency
Research participants described fatigue, frustration, and difficulty with controller-based interaction, making hands-free session control a functional accessibility requirement rather than a convenience layer.
•Recognition Failure Could Break Trust
Early trials exposed false triggers, response delay, command recall, and accent support as barriers, so listening state, confirmation, error recovery, and fallback guidance had to remain visible.
Controller-access problem frame
The thesis brief frames precise controller input and complex gesture navigation as barriers for people with limited upper-limb mobility and low muscle tone.
02 / Mandate
Mandate
•Make Voice the Primary Control Channel
Support session start, progression, completion, breathing, visualization, lighting, and affirmation interactions without requiring a handheld controller.
•Make System State Perceivable
Pair spoken feedback with gaze-anchored text, color, sound, and confidence cues so users could tell when the system was listening, processing, succeeding, or asking for clarification.
•Measure Accessibility Through Iteration
Use prototype metrics and participant sessions to evaluate task completion, usability, latency, comfort, recognition limits, and the ten-pillar interaction framework.
•Keep Incomplete Work Explicit
Separate implemented features from partial and planned pillars, especially paraphrase coverage, pacing, redirection, privacy controls, and broader language support.
Ten-pillar implementation ledger
The final framework records six implemented pillars, two partial pillars, and two planned pillars rather than presenting the accessibility system as complete.
03 / Build
Build
1.Grounded the System in Participatory Research
Combined literature review, expert input, a 42-person survey, participatory co-design, and accessibility synthesis to translate lived barriers into hands-free agency, transparent feedback, and flexible pacing requirements.
2.Iterated Across Three Prototype Generations
Moved from a Wizard-of-Oz paper and Figma test to a live Quest 3 alpha in Unity, then to a beta with faster response, confidence cues, adaptive lighting, and positive affirmation modules.
3.Implemented the Quest 3 Voice System
Built the Unity 2022.3 application with Meta Voice SDK and Wit.ai, wake-word activation, command variants, intent and confidence handling, TTS, audio and visual feedback, world-space UI, and voice-controlled therapy modules.
4.Tested with Eight Final Participants
Ran structured Quest 3 sessions across varied ages, XR familiarity, ability, and professional backgrounds, measuring SUS, voice-only task completion, response latency, error recovery, comfort, and feedback clarity.
Research-to-prototype record
The build record connects literature review and co-design with ShapesXR exploration, Unity development, voice recognition, multimodal feedback, and structured participant testing.
04 / Outcomes
Outcomes
•Measured Alpha-to-Beta Improvement
The process-book record reports latency improving from 530 ms to 380 ms, voice-only task completion increasing from 64% to 79%, and SUS increasing from 71 to 79 between the alpha and beta prototypes.
•Six Pillars Implemented
The final framework records six of ten accessibility pillars as implemented, two as partial, and two as planned; this status is preserved instead of presenting the system as universally complete.
•Known Recognition Limits
Participant sessions identified English-only recognition, accent sensitivity, command recall, and confidence-cue clarity as remaining constraints, providing a concrete research agenda rather than an unsupported scalability claim.
Documented participant outcomes
The result record summarizes reported comfort and engagement gains while keeping broader scalability as a direction for future research.
Exhibition & Future Directions
•Public Thesis Exhibition
Delivered a gallery installation with a Quest headset, live Unity experience, project framing, and process evidence for hands-on public review.
•Thesis Defense and Process Book
Presented the research, prototype progression, participant findings, ten-pillar framework, and limitations in a public M.Des defense and a 41-page process book.
•Defined Follow-Up Work
Broader accent and language support, clearer redirection, privacy controls, on-device storage choices, and deeper personalization remain explicit future work.
Interactive thesis exhibition
The completed gallery installation pairs the Quest headset, live Unity environment, project framing, and process evidence in a public hands-on review setting.
Exhibition plan
The planning artifact defines the prototype station, participant feedback, and future gesture and personalization concepts before the final installation was delivered.
Continue the evidence trail
From proof to role fit
Compare SpeakEasy with adjacent systems, or carry its reviewed capabilities into an Adaptive Focus brief.