Voice-first AR systems / Cultural portal navigation
Portals
Led SnapAR development for a four-person Spectacles hackathon prototype that opens and routes cultural portals through VoiceML commands, supports explicit return phrases, and lets users resize portal objects with hand-tracked pinch input.
- Operating proof
- Direct hackathon prototype evidence for voice-first AR navigation, spanning VoiceML portal activation, phrase-routed destination states, explicit return commands, hand-tracked pinch scaling, spatial media assets, and an on-device Spectacles scene.
- Engagement
- Stanford Immerse the Bay 2024 / Four-person team
- Evidence set
- 5 reviewed artifacts
- Capability coverage
- 4 documented areas
Primary artifact
01 / 05Three-destination Spectacles scene
The primary on-device artifact places Samoa, India, and San Francisco portals in one hackathon scene with visible spoken-entry prompts instead of a conventional menu.
01 / Context
Situation
•A four-person team had one hackathon to make culture spatial
At Stanford's Immerse the Bay 2024 XR Hackathon, the team explored how Snap Spectacles could turn cultural destinations into portals combining place, music, voice, and embodied interaction. Mike served as project lead and SnapAR developer alongside cultural research, accessibility, audio, UX, and Lens Studio integration roles.
•The interface needed to stay out of the view
A portal experience loses its effect if every action requires a dense visual menu. The prototype needed direct spoken commands for entering, leaving, activating, and deactivating destinations while preserving a hand-based manipulation path.
•The concept was larger than the implementation window
The broader pitch included cultural timelines, collaborative music, computer-vision recognition, educational layers, and climate awareness. The repository's clearest implemented proof is the voice-and-hand portal interaction system, so the case study separates that core from the roadmap.
Four-person hackathon showcase
The submitted Devpost frame preserves the project identity and four-person team context around the voice-and-hand Spectacles build.
02 / Mandate
Mandate
•Represent three destinations in one scene
Build distinct portal states for Uttar Pradesh, India; Afu Aau Falls, Samoa; and San Francisco, USA, with destination-specific objects and media that could be selected without conventional navigation chrome.
•Make activation and exit language explicit
Support simple phrases for turning the portal system on or off, entering a destination, leaving the portal, and returning home, with final-transcription checks before scene state changes.
•Keep a second embodied input path
Add hand-tracked pinch scaling so portal objects could be manipulated directly rather than making voice the only available interaction mode.
•Package a judgeable prototype
Connect the Lens Studio scene, portal model, VoiceML module, hand-tracking package, cultural audio and video assets, public repository, and Devpost demonstration into one hackathon submission.
03 / Build
Build
1.Built a guarded VoiceML portal lifecycle
Implemented portal-on, activate-portal, portal-off, and deactivate-portal keyword groups; validated required scene inputs; prevented competing VoiceML sessions; forced a disabled initial state; and handled listening updates, failures, activation, animation, and deactivation.
2.Routed spoken destinations into scene states
Mapped enter India, enter Samoa, enter America, exit portal, leave portal, and return home phrases to three destination objects and a controlled initial state, processing only final English transcriptions.
3.Added hand-tracked portal scaling
Connected Lens Studio's hand-tracking component to pinch detection, captured the starting pinch distance and object scale, and applied uniform scaling while the gesture remained active.
4.Integrated cultural media into the Lens scene
Assembled the Holi portal model, Holi instrumental and spoken-description media, Billie Holiday audio, device-camera texture, hand-tracking controller, VoiceML module, and Spectacles Interaction Kit inside the submitted Lens Studio project.
Guarded VoiceML portal lifecycle
The implemented Lens Studio script validates scene inputs, prevents competing VoiceML sessions, starts from a disabled portal state, and handles activation, animation, deactivation, and listening errors.
Phrase-routed destination states
The routing source defines explicit phrases for entering India, Samoa, and America and for exiting, leaving, or returning home, then maps final transcriptions into controlled scene-object states.
04 / Outcomes
Outcomes
•Three destinations became phrase-addressable
The implemented command vocabulary routes India, Samoa, and America portal states and provides three explicit return phrases, making the scene operable without an always-visible destination menu.
•Voice and hand input shared the interaction model
VoiceML controls scene state while tracked pinch distance controls portal scale. The prototype demonstrates two complementary XR inputs rather than claiming that voice alone resolves every accessibility need.
•The work remained inspectable after the event
The Lens Studio scene, custom JavaScript, spatial media assets, team roles, demo video, Devpost narrative, and MIT-licensed repository preserve the hackathon build as reviewable implementation evidence.
•A bounded prototype, not a completed cultural platform
The repository proves portal activation, destination routing, pinch scaling, and packaged media. Computer-vision recognition, cultural timelines, global jam sessions, sharing, environmental-data triggers, and broader accessibility validation remain concept or future work.
Hand-tracked pinch scaling
A second input path reads active hand and pinch state, captures initial distance and scale, and uniformly resizes the portal while the gesture remains active.
Continue the evidence trail
From proof to role fit
Compare Portals with adjacent systems, or carry its reviewed capabilities into an Adaptive Focus brief.