Back to systems
CASE FILE / AF-14REVIEWED EVIDENCEACTIVE SYSTEM / November 2024

Voice-first AR systems / Cultural portal navigation

Portals

Led SnapAR development for a four-person Spectacles hackathon prototype that opens and routes cultural portals through VoiceML commands, supports explicit return phrases, and lets users resize portal objects with hand-tracked pinch input.

Operating proof
Direct hackathon prototype evidence for voice-first AR navigation, spanning VoiceML portal activation, phrase-routed destination states, explicit return commands, hand-tracked pinch scaling, spatial media assets, and an on-device Spectacles scene.
Engagement
Stanford Immerse the Bay 2024 / Four-person team
Evidence set
5 reviewed artifacts
Capability coverage
4 documented areas
SnapAR PrototypingVoice Interaction DesignSpatial Experience DesignHackathon Product Leadership

Primary artifact

01 / 05
artifact_viewer.sh

Three-destination Spectacles scene

The primary on-device artifact places Samoa, India, and San Francisco portals in one hackathon scene with visible spoken-entry prompts instead of a conventional menu.

01 / Context

Situation

A four-person team had one hackathon to make culture spatial

At Stanford's Immerse the Bay 2024 XR Hackathon, the team explored how Snap Spectacles could turn cultural destinations into portals combining place, music, voice, and embodied interaction. Mike served as project lead and SnapAR developer alongside cultural research, accessibility, audio, UX, and Lens Studio integration roles.

The interface needed to stay out of the view

A portal experience loses its effect if every action requires a dense visual menu. The prototype needed direct spoken commands for entering, leaving, activating, and deactivating destinations while preserving a hand-based manipulation path.

The concept was larger than the implementation window

The broader pitch included cultural timelines, collaborative music, computer-vision recognition, educational layers, and climate awareness. The repository's clearest implemented proof is the voice-and-hand portal interaction system, so the case study separates that core from the roadmap.

evidence_situation.log

Four-person hackathon showcase

The submitted Devpost frame preserves the project identity and four-person team context around the voice-and-hand Spectacles build.

02 / Mandate

Mandate

Represent three destinations in one scene

Build distinct portal states for Uttar Pradesh, India; Afu Aau Falls, Samoa; and San Francisco, USA, with destination-specific objects and media that could be selected without conventional navigation chrome.

Make activation and exit language explicit

Support simple phrases for turning the portal system on or off, entering a destination, leaving the portal, and returning home, with final-transcription checks before scene state changes.

Keep a second embodied input path

Add hand-tracked pinch scaling so portal objects could be manipulated directly rather than making voice the only available interaction mode.

Package a judgeable prototype

Connect the Lens Studio scene, portal model, VoiceML module, hand-tracking package, cultural audio and video assets, public repository, and Devpost demonstration into one hackathon submission.

03 / Build

Build

1.Built a guarded VoiceML portal lifecycle

Implemented portal-on, activate-portal, portal-off, and deactivate-portal keyword groups; validated required scene inputs; prevented competing VoiceML sessions; forced a disabled initial state; and handled listening updates, failures, activation, animation, and deactivation.

2.Routed spoken destinations into scene states

Mapped enter India, enter Samoa, enter America, exit portal, leave portal, and return home phrases to three destination objects and a controlled initial state, processing only final English transcriptions.

3.Added hand-tracked portal scaling

Connected Lens Studio's hand-tracking component to pinch detection, captured the starting pinch distance and object scale, and applied uniform scaling while the gesture remained active.

4.Integrated cultural media into the Lens scene

Assembled the Holi portal model, Holi instrumental and spoken-description media, Billie Holiday audio, device-camera texture, hand-tracking controller, VoiceML module, and Spectacles Interaction Kit inside the submitted Lens Studio project.

evidence_action.log

Guarded VoiceML portal lifecycle

The implemented Lens Studio script validates scene inputs, prevents competing VoiceML sessions, starts from a disabled portal state, and handles activation, animation, deactivation, and listening errors.

Phrase-routed destination states

The routing source defines explicit phrases for entering India, Samoa, and America and for exiting, leaving, or returning home, then maps final transcriptions into controlled scene-object states.

04 / Outcomes

Outcomes

Three destinations became phrase-addressable

The implemented command vocabulary routes India, Samoa, and America portal states and provides three explicit return phrases, making the scene operable without an always-visible destination menu.

Voice and hand input shared the interaction model

VoiceML controls scene state while tracked pinch distance controls portal scale. The prototype demonstrates two complementary XR inputs rather than claiming that voice alone resolves every accessibility need.

The work remained inspectable after the event

The Lens Studio scene, custom JavaScript, spatial media assets, team roles, demo video, Devpost narrative, and MIT-licensed repository preserve the hackathon build as reviewable implementation evidence.

A bounded prototype, not a completed cultural platform

The repository proves portal activation, destination routing, pinch scaling, and packaged media. Computer-vision recognition, cultural timelines, global jam sessions, sharing, environmental-data triggers, and broader accessibility validation remain concept or future work.

evidence_result.log

Hand-tracked pinch scaling

A second input path reads active hand and pinch state, captures initial distance and scale, and uniformly resizes the portal while the gesture remains active.

Continue the evidence trail

From proof to role fit

Compare Portals with adjacent systems, or carry its reviewed capabilities into an Adaptive Focus brief.

XR and spatial computingVoice interactionAccessibility