Calls and questions repeat all day
Orders, opening hours, prices and bookings follow the same pattern, and staff answer them between other work.
AI voice agents built on Pipecat and Gemini Live. Real-time voice assistants that answer and act through your own tools, with scoped rules, confirmed actions and a handoff to your team.
We build real-time voice agents that hold a natural conversation and act through your own systems: look up a menu, check a price, place an order, book a call, or pass the conversation to a person. They run on Pipecat with Gemini Live, the same stack behind our OpenPOS phone-ordering agent and the KUKI cooking assistant.
A good fit for
Businesses that answer the same questions or take the same orders by voice, many times a day.
What we work toward
A voice agent with a clear job, connected tools, written rules and a tested handoff to your team.
Orders, opening hours, prices and bookings follow the same pattern, and staff answer them between other work.
Some customers would rather speak, and some moments are hands-free: in the kitchen, on the move, or on the phone.
A voice agent that guesses prices or availability creates problems. It needs tools that read your system, and rules for what it may promise.
The final scope follows your use case and systems. These are the parts we build and connect.
Pipecat with Gemini Live speech-to-speech, tuned turn-taking so the agent waits for natural pauses, and long-session handling.
Function calls into your APIs, such as menu search, cart and order placement in OpenPOS, or knowledge search and packages on this website.
What the agent may say and promise, which actions need the caller's confirmation, and how it hands over to a person.
A WebRTC voice client for your site or test console, Docker deployment with the networking WebRTC needs, and idle and session limits for cost control.
Pick one conversation the agent owns, the data it needs and the actions it may take.
Build the function calls into your systems so answers and prices come from your systems rather than the model.
Persona, languages, confirmation steps and the handoff path, tested against real conversations in the browser.
Ship the agent with monitoring, review conversations and outcomes, and adjust prompts and tools from what callers say.
Our voice work started with real products. For OpenPOS we built a phone-ordering agent. At the start of each call it loads the branch's live menu from the POS. It searches the menu, builds the cart, sets takeaway or delivery and records the customer's name and number. It then reads the whole order and total back, and only after a clear yes places the order into the POS, where it prints to the kitchen like any other order. The POS calculates every total.
KUKI is a hands-free cooking assistant on the same stack. It reads each recipe step aloud and waits until the cook says they are done. It runs timers on the server, announces when they finish, scales portions and suggests substitutions, in Swedish or English. Running it in real kitchens taught us how to tune turn-taking for people who pause mid-sentence, and how to keep long sessions alive.
The assistant on this website is the third. It uses the same knowledge and tools as our chat assistant, speaks English, Urdu or Hindi and quotes prices only from our published packages. It sends you to the booking page instead of booking by voice.
WebRTC needs STUN and a published UDP range to connect from real networks. We run agents in Docker with those settings, plus idle cut-offs and session limits so cost stays predictable.
Prices, availability and facts come from your systems. When the systems have no answer, the agent says it does not know.
The agent reads back orders and other changes and waits for an explicit yes.
Callback requests, contact details or a booking link whenever the agent cannot help.
Tested WebRTC connectivity, idle cut-off and a maximum session length.
Starter
$999/ project
One voice agent for one job, tested in the browser on real cases.
Growth
$2,499+ $299 / month
A deployed agent that writes into your system and hands off to people.
Scale
From $4,999+ $599 / month
Several agents, channels or locations on one setup.
Prices cover the build. The model provider bills usage at cost per minute of audio. Production includes monitoring, conversation review and fixes for $299/month.
We have built three. The first is a phone-ordering agent for OpenPOS restaurants that reads the branch's live menu, builds the order, reads it back and places it into the POS. The second is KUKI, a hands-free cooking assistant that reads recipe steps aloud, runs timers and scales portions. The third is the voice assistant on this website, which shares the chat assistant's knowledge and tools.
Pipecat, an open-source framework for real-time voice pipelines, with Google's Gemini Live model for speech-to-speech conversation and function calling. Audio travels over WebRTC, so a visitor can talk from the browser without installing an app.
The agent replies in the language the caller speaks. The OpenPOS agent is set up to follow the caller across major languages such as English, Arabic, Turkish and Urdu, with a fallback when a language is not supported. KUKI runs in Swedish and English. We agree the languages for your agent during scoping and test them.
Yes, through tools. Prices come from your system rather than the model. In the OpenPOS agent, the POS backend calculates every total. Before placing an order, the agent reads it back and needs an explicit yes from the caller.
It hands off. The OpenPOS agent records a callback request with the caller's details for staff, and this website's assistant shows WhatsApp, phone and email details or opens the booking page. We define the handoff rules for your agent before launch.
A pilot for one use case is $999. A production agent deployed into your system is $2,499 plus $299/month for monitoring and improvements. Several agents, channels or locations start at $4,999 plus $599/month. Running costs include the voice model, which the provider bills per minute of audio, and hosting.
More in AI & automation
A free 15-minute call, then a written scope and estimate. No obligation, no sales script.
Talk to Wasevo AI
Voice · English, Urdu, Hindi
Try the voice agent
Ask about our services, prices, projects or booking a call. Your microphone starts only when you press Start.