How AI agents redefine universal design to improve accessibility

Machine Learning


Our research direction: Designing for accessibility

Our early research found that a significant barrier to digital equity is the “accessibility gap,” or the delay between releasing a new feature and creating a layer of support for it. To fill this gap, we are moving from reactive tools to agent systems that are native to the interface.

Research pillar: Improving accessibility using multisystem agents

Multimodal AI tools offer one of the most promising avenues for building accessible interfaces. Certain prototypes, such as our web readability efforts, tested models where a central orchestrator acts as a strategic reading manager.

Instead of users having to navigate a complex maze of menus, Orchestrator maintains a shared context and delegates tasks to specialized subagents to make documents easier to understand and access.

  • Summary agent: Master complex documents and make even the deepest insights clearly accessible by breaking down information and delegating key tasks to specialized sub-agents.
  • Configuration agent: Dynamically handles UI adjustments such as text scaling.

By testing this modular approach, our research has shown that users can navigate the system more intuitively, users don’t have to hunt for the “right” button, and specialized tasks are always handled by the right expert.

Aiming for multimodal fluency

Our research also focuses on moving beyond basic text-to-speech to multimodal fluency. By leveraging Gemini’s ability to process audio, visual, and text simultaneously, we built a prototype that can instantly transform live video into interactive audio descriptions.

This is more than just describing a scene. It’s about situational awareness. In our co-design sessions, we observed how allowing users to interactively query the environment, asking for specific visual details as they occur, can reduce cognitive load and transform passive experiences into active conversational exploration.



Source link