Multimodal Agents by Sierra are AI agents that bring voice, text, and visuals into the same customer conversation. Sierra has long believed that the conversation is the interface: the customer says what they need and the agent figures out the rest. Multimodal agents extend that belief beyond a single medium. Rather than making customers choose between talking, typing, or looking at something, the agent gives them the best of each — voice to explain what you need, a visual to compare options side by side, and text when you want to reference something later. The result is a single, continuous conversation that adapts to what the customer is trying to accomplish at that moment.
The problem this solves is familiar to anyone who has tried to make a decision over the phone. Sierra describes trying to upgrade a mobile plan over the phone: the representative talks through models, colors, storage sizes, and monthly rates, and the customer is left comparing things in their head and picking a phone they cannot picture. Voice is genuinely good for parts of that interaction — you can say what you actually need and ask questions more easily than you can over text — but you cannot see the thing you are about to buy. Multimodal agents close that gap. Stitching channels together is not the hard part; the real trick, as Sierra puts it, is knowing which one to use when.
That is where the agent's judgement comes in. Agents built on Sierra anticipate what is needed for each conversation and automatically shift between modes — voice, visuals, or text — without making the customer start over or repeat themselves. The choice of medium follows the shape of the task: voice to explain what you need, a visual to compare options side by side, or text when you want to reference something later. Because that switching is automatic, the customer never has to manage the interface. They simply continue the conversation, and the agent keeps the relevant context intact as the medium changes, which is precisely what prevents the restarting and repeating that usually happens when a support interaction jumps between a phone call, a chat window, and a web page.
Visuals are part of the conversation rather than a separate destination. Sierra's example is a disrupted flight: you call the airline to get a new flight, and instead of a representative reading off alternate options one by one, you see them laid out with departure times, layovers, and pricing right in the conversation. You pick one, and the agent keeps going from there. Choosing a seat works the same way — you see the map and tap the seat you want. And for times when it is easier to talk than type, you can switch to voice and explain exactly what you need; the agent captures those details without making you type a paragraph into a text box. In each case the visual carries the comparison work that language handles poorly, while voice carries the nuance that menus and forms handle poorly.
Sierra's approach is also designed to avoid rebuilding the same experience for every place the agent lives. With Sierra, you can build your agent once and easily deploy across all channels, and the same is true for multimodal agents: once you build a visual component, your agent can use it everywhere it lives. A comparison table or a calendar does not need to be recreated for each surface the agent operates on. That single-build approach reduces duplication for the team maintaining the experience and keeps behaviour consistent for the customer, who encounters the same kind of interactive element regardless of where the conversation happens to be taking place.
Sierra's MCP UI integration is what lets teams bring interactive components into the conversation: product cards, comparison tables, calendars, and forms. Those components are designed and hosted by your own team, so you decide how they look, what they show, and when they change. That control matters because the visual layer is often the part of a customer experience that carries brand and merchandising decisions, not just function. Because your team hosts the components, when you make an update it is automatically reflected everywhere without needing to redeploy or maintain different versions for each platform. And when a component needs more room, it can expand to full screen to show calendars, long comparison tables, multi-step forms, and more — so the same building block can serve as an inline detail inside a conversation or as a focused, full-attention task when the customer needs to complete something substantial.
Taken together, the methodology is straightforward: keep the conversation as the interface, let the agent decide which medium each moment calls for, and make the interactive pieces reusable across every surface. Sierra frames the goal as customers never having to choose. On one call, customers can talk through what they need, glance at a screen to compare their options, and tap to confirm — without ever pausing the conversation to switch tools. The agent, not the customer, manages the transitions, which is what makes an interaction that spans voice, visuals, and text feel like a single continuous exchange rather than three separate ones.
The outcomes described are practical. Customers get through decisions faster because they can see options while hearing about them, and they avoid the frustration of describing the same need twice or rebuilding context after a channel change. They can also reference something later in text when that is easier than listening. Businesses, meanwhile, get a single deployment path: build the agent and its visual components once, use them across channels, update them in one place, and avoid maintaining separate versions per platform. And the experience is described as being as easy to build and deploy as it is for customers to use, which lowers the practical barrier to offering a multimodal customer experience at all.
Concrete workflows in Sierra's own examples include upgrading a mobile plan, where a customer talks through what they need and compares phones, colors, storage sizes, and monthly rates visually instead of holding the options in their head. A disrupted flight is another: the agent surfaces alternate flights with departure times, layovers, and pricing in the conversation, and the customer picks one and continues. Seat selection follows the same pattern, with a map the customer taps rather than a description they have to parse. Voice-first moments are covered too — when it is easier to explain something than to type it, the customer can switch to voice and the agent captures the details. More broadly, any conversation that involves comparing options side by side, filling in a form, or choosing a time can use interactive components inside the exchange itself.
Multimodal Agents by Sierra are aimed at organizations that handle customer conversations and want those conversations to adapt to the customer rather than the other way around — customer experience and support functions in particular. The people who build the visual layer are the customer's own teams: Sierra states that your team designs and hosts the components used in the conversation. Deployment is described in terms of channels rather than a single app, since the same agent and the same visual components are meant to work everywhere the agent lives. Sierra's MCP UI integration is the mechanism named in the content for bringing interactive components such as product cards, comparison tables, calendars, and forms directly into a conversation.
The core idea behind Multimodal Agents by Sierra is that the best interface is the one the conversation needs. Voice, visuals, and text stop being competing options and become modes the agent moves between as the situation changes — voice when explaining is easier, a visual when comparing side by side helps, text when something needs to be referenced later. Because agents built on Sierra anticipate what is needed and shift automatically, customers never start over or repeat themselves, and because visual components are built once and hosted by your team, they can appear everywhere the agent works. That is the promise: one agent, every surface, and a conversation that morphs to fit the customer.