Book a free call
AI & automation

AI voice agents

AI voice agents built on Pipecat and Gemini Live. Real-time voice assistants that answer and act through your own tools, with scoped rules, confirmed actions and a handoff to your team.

Gemini LiveSpeech to speech
PipecatOpen-source voice pipeline
WebRTCBrowser audio, no app
Project, then monthlyBest for: Teams that take orders or questions by voice
the opportunity

A voice agent that takes actions in your systems.

We build real-time voice agents that hold a natural conversation and act through your own systems: look up a menu, check a price, place an order, book a call, or pass the conversation to a person. They run on Pipecat with Gemini Live, the same stack behind our OpenPOS phone-ordering agent and the KUKI cooking assistant.

A good fit for

Businesses that answer the same questions or take the same orders by voice, many times a day.

Restaurant phone ordersWebsite assistantsHands-free product guidesEnquiries and bookingMultilingual customers

What we work toward

A voice agent with a clear job, connected tools, written rules and a tested handoff to your team.

recognize the problem

When a voice agent is worth building.

01 / 03

Calls and questions repeat all day

Orders, opening hours, prices and bookings follow the same pattern, and staff answer them between other work.

02 / 03

A chatbot alone does not fit

Some customers would rather speak, and some moments are hands-free: in the kitchen, on the move, or on the phone.

03 / 03

Answers must match your real data

A voice agent that guesses prices or availability creates problems. It needs tools that read your system, and rules for what it may promise.

what is included

What we build.

The final scope follows your use case and systems. These are the parts we build and connect.

01

Real-time voice pipeline

Pipecat with Gemini Live speech-to-speech, tuned turn-taking so the agent waits for natural pauses, and long-session handling.

02

Tools on your data

Function calls into your APIs, such as menu search, cart and order placement in OpenPOS, or knowledge search and packages on this website.

03

Rules and handoff

What the agent may say and promise, which actions need the caller's confirmation, and how it hands over to a person.

04

Browser client and deployment

A WebRTC voice client for your site or test console, Docker deployment with the networking WebRTC needs, and idle and session limits for cost control.

a practical workflow

From one clear job to a live agent.

  1. 01

    Define the job

    Pick one conversation the agent owns, the data it needs and the actions it may take.

  2. 02

    Connect the tools

    Build the function calls into your systems so answers and prices come from your systems rather than the model.

  3. 03

    Write the rules

    Persona, languages, confirmation steps and the handoff path, tested against real conversations in the browser.

  4. 04

    Deploy and review

    Ship the agent with monitoring, review conversations and outcomes, and adjust prompts and tools from what callers say.

how wasevo helps

Built from agents that are already working.

Our voice work started with real products. For OpenPOS we built a phone-ordering agent. At the start of each call it loads the branch's live menu from the POS. It searches the menu, builds the cart, sets takeaway or delivery and records the customer's name and number. It then reads the whole order and total back, and only after a clear yes places the order into the POS, where it prints to the kitchen like any other order. The POS calculates every total.

KUKI is a hands-free cooking assistant on the same stack. It reads each recipe step aloud and waits until the cook says they are done. It runs timers on the server, announces when they finish, scales portions and suggests substitutions, in Swedish or English. Running it in real kitchens taught us how to tune turn-taking for people who pause mid-sentence, and how to keep long sessions alive.

The assistant on this website is the third. It uses the same knowledge and tools as our chat assistant, speaks English, Urdu or Hindi and quotes prices only from our published packages. It sends you to the booking page instead of booking by voice.

WebRTC needs STUN and a published UDP range to connect from real networks. We run agents in Docker with those settings, plus idle cut-offs and session limits so cost stays predictable.

delivery checks

What we check before launch.

Answers from your systems

Prices, availability and facts come from your systems. When the systems have no answer, the agent says it does not know.

Confirmation before action

The agent reads back orders and other changes and waits for an explicit yes.

A working handoff

Callback requests, contact details or a booking link whenever the agent cannot help.

Connection and cost limits

Tested WebRTC connectivity, idle cut-off and a maximum session length.

How we work
Answers only from your data and tools
Prices come from your system, never the model
The caller confirms before anything is placed
Handoff to a person at any point
Compare packages

Pick the scope that fits

Starter

Pilot

$999/ project

One voice agent for one job, tested in the browser on real cases.

  • One use case, such as orders or enquiries
  • Tools connected to your own data
  • Rules for prices, promises and handoff
  • Browser test console
  • Review of real conversations
Book a call

Growth

Production

$2,499+ $299 / month

A deployed agent that writes into your system and hands off to people.

  • Everything in Pilot
  • Actions confirmed by the caller first
  • Handoff or callback requests
  • Docker deployment with WebRTC networking
  • Cost controls: idle cut-off and session limits
Book a call

Scale

Scale

From $4,999+ $599 / month

Several agents, channels or locations on one setup.

  • Agents for several use cases or branches
  • More languages, tested per agent
  • Conversation review and improvements
  • Written scope first
Talk to us

Prices cover the build. The model provider bills usage at cost per minute of audio. Production includes monitoring, conversation review and fixes for $299/month.

FAQ

AI voice agents: questions answered

Can't see yours? Ask on WhatsApp

What have you built with voice agents?

We have built three. The first is a phone-ordering agent for OpenPOS restaurants that reads the branch's live menu, builds the order, reads it back and places it into the POS. The second is KUKI, a hands-free cooking assistant that reads recipe steps aloud, runs timers and scales portions. The third is the voice assistant on this website, which shares the chat assistant's knowledge and tools.

Which technology do you use?

Pipecat, an open-source framework for real-time voice pipelines, with Google's Gemini Live model for speech-to-speech conversation and function calling. Audio travels over WebRTC, so a visitor can talk from the browser without installing an app.

Which languages can the agent speak?

The agent replies in the language the caller speaks. The OpenPOS agent is set up to follow the caller across major languages such as English, Arabic, Turkish and Urdu, with a fallback when a language is not supported. KUKI runs in Swedish and English. We agree the languages for your agent during scoping and test them.

Can the agent quote prices or take orders?

Yes, through tools. Prices come from your system rather than the model. In the OpenPOS agent, the POS backend calculates every total. Before placing an order, the agent reads it back and needs an explicit yes from the caller.

What happens when the agent cannot help?

It hands off. The OpenPOS agent records a callback request with the caller's details for staff, and this website's assistant shows WhatsApp, phone and email details or opens the booking page. We define the handoff rules for your agent before launch.

How much does a voice agent cost?

A pilot for one use case is $999. A production agent deployed into your system is $2,499 plus $299/month for monitoring and improvements. Several agents, channels or locations start at $4,999 plus $599/month. Running costs include the voice model, which the provider bills per minute of audio, and hosting.

Ready to start with AI voice agents? Book a free call

A free 15-minute call, then a written scope and estimate. No obligation, no sales script.