AI voice agents now handle calls for reception, support, bookings and sales. But between a demo that sounds good and a system that works in production, there are 7 technical decisions that determine whether you will save time or inherit a problem. This article breaks down the criteria that matter when you evaluate providers: latency, voices, integrations, languages, scalability, privacy and real cost.
If you are comparing platforms or have received a proposal and do not know what to ask, here is the checklist that technical teams use before signing.
This page is for information only and is not binding advice. Each case is tailored after a diagnosis.
TL;DR
- Latency below 1.5 seconds is critical; above 2 seconds the experience degrades.
- High-end synthetic voices (ElevenLabs, OpenAI) sound natural; clone a voice only if your brand identity justifies it.
- Integrate the agent with your CRM, switchboard and knowledge base via webhook or REST API; without integrations it is a silo.
- The agent must handle your language and regional variants (Castilian Spanish, Catalan, Galician) with transcription accuracy above 95%.
- Scalability: the system must withstand 10x peaks without degrading latency or voice quality.
- Privacy and RGPD (the Spanish name for the EU’s GDPR) compliance require encrypted recordings, configurable TTL and EU data residency.
- Typical market cost: €0.10-0.40/minute + platform fee; custom solutions from €800/month.
Latency: the invisible bottleneck
Latency is the time between the customer finishing speaking and the agent starting to respond. A natural conversation requires less than 1.5 seconds; above 2 seconds it feels robotic and increases drop-off.
The full chain includes:
- Speech-to-text transcription (STT): 200-500 ms with Whisper or Deepgram.
- LLM processing (response): 400-1200 ms depending on model and context.
- Text-to-speech synthesis (TTS): 300-800 ms with neural voices.
- Network and jitter: 100-300 ms depending on telephony provider.
A good provider measures P95 latency (95th percentile) and shows it to you on a dashboard. If there is no public metric, ask for it. If the demo does not mention it, that is a warning sign.
How to verify latency in tests
- Make 10 test calls at different times of day.
- Record your experience and measure the silence with a stopwatch.
- If it exceeds 2 seconds in 3 out of 10 calls, rule the provider out.
Latency is not fixed by “future optimisations”. It is architecture. A well-designed AI phone agent starts with latency < 1.2 s at P95.
Natural voices vs cloned voices
The synthetic voices of 2026 (ElevenLabs Turbo v2.5, OpenAI Alloy, Azure Neural) sound human, handle intonation and respect pauses. Cloned voices of your team bring identity but add complexity.
When to use synthetic voices
- You start quickly, with no need to capture samples.
- You need multiple languages or accents.
- The cost per minute is predictable (€0.05-0.15).
When to clone a voice
- Your brand has a strong vocal identity (radio, podcast, TV).
- The customer recognises the voice of the founder or spokesperson.
- You have 30-60 minutes of clean audio and the speaker’s RGPD consent.
Cloning costs between €200 and €800 setup depending on the provider, plus an additional €0.10-0.20 per minute of synthesis. If you do not have a clear business case, start with synthetic voices.
Integrations: CRM, switchboard and knowledge base
An isolated agent is a dead end. It must read data from your CRM, write results back and transfer calls to humans when it cannot resolve them.
Basic integrations
- CRM (Salesforce, HubSpot, Pipedrive): the agent reads the customer record from the incoming phone number and writes a post-call note.
- Switchboard (Asterisk, 3CX, Twilio): warm or cold transfer to a human extension.
- Knowledge base (Notion, Confluence, Google Drive): the agent looks up answers via RAG before responding.
How to verify integration capability
- Ask for the API or webhook documentation.
- Ask for examples of JSON payloads.
- Request a use case of integration with your current stack.
If the provider has no documented API or only offers “native integrations” with 3 tools, flexibility is nil. A serious platform exposes an HTTP POST webhook on every call event. That way your team can connect any system via automation.
Languages and regional variants
STT models trained on English fail with Andalusian Spanish, Catalan or Galician. Transcription accuracy must exceed 95% in your target language and accent.
Language checklist
- Castilian Spanish from Spain: check that it understands “vale”, “anda”, “guay”, “ostras”.
- Catalan: if your business operates in Catalonia, insist on a test with native speakers.
- Sector vocabulary: medicine, law and property have specific jargon. The model must be fine-tuned or given a custom glossary.
Some providers offer STT fine-tuning for €500-2000. If your business uses technical terms (product codes, own brands, acronyms), fine-tuning cuts errors from 15% to 3%.
Scalability: from 10 to 1,000 calls a day
An agent that works with 50 calls/day can collapse on Black Friday or during a marketing campaign. Scalability is not just adding servers; it is stable latency under load.
Stress scenarios
- 10x peak: your business receives 500 calls in 2 hours. The agent must keep latency < 1.5 s and 0% dropped calls.
- Simultaneous calls: 20 customers call at once. The provider must have an instance pool or autoscaling.
Questions for the provider
- What is the limit of simultaneous calls on your plan?
- How does the system scale at peaks? (autoscaling, reserved instances, queue).
- What happens if I exceed the limit? (queue, busy tone, dropped call).
Serious SaaS providers publish an availability SLA (99.5% or higher) and P95 latency. If there is no written SLA, the system is not ready for production.
Privacy and RGPD compliance
Calls contain personal data. The agent must comply with RGPD: consent, encryption, EU data residency and recording TTL.
Minimum requirements
- Consent: the agent states at the start that the call is being recorded.
- Encryption: audio and transcripts encrypted at rest (AES-256) and in transit (TLS 1.3).
- Data residency: servers in the EU (Frankfurt, Amsterdam, Paris). Not AWS us-east.
- TTL: recordings are deleted after 30, 60 or 90 days according to the retention policy.
If the provider does not mention RGPD or says “we comply with GDPR” without documentation, ask for a DPA (Data Processing Agreement) and a technical annex. A serious provider has a template DPA ready to sign.
For more legal context, see AI and data protection.
Real cost: hidden variables
The public prices of SaaS providers tend to omit variable costs. A business with 200 calls of 3 minutes a month can pay between €110 and €400 depending on provider and plan.
Typical market cost structure
- Base fee: €50-200/month (platform, dashboard, support).
- Call minutes: €0.10-0.40/minute.
- Transcription: included or €0.006/minute extra.
- Cloned voice: +€0.10/minute if you use a custom voice.
- Premium LLM: +€0.05/minute if you use GPT-4 instead of a base model.
Example of an indicative monthly cost
Business with 200 calls/month averaging 3 minutes:
- 600 minutes × €0.20 = €120 (calls).
- €100 (base fee).
- Total: €220/month.
If volume grows to 1,000 calls (3,000 min), the cost can jump to €700/month. Custom solutions with their own infrastructure start from €800/month but cut the variable cost to €0.05-0.10/minute once the initial investment has been paid off.
STAKKER works with a free diagnosis and a tailored proposal; each project is adjusted to the client’s volume and stack. If you need a cost projection, get in touch here.
Transfer to humans: the necessary plan B
The agent must recognise when it cannot resolve an issue and transfer the call to a human without the customer having to repeat their problem.
Types of transfer
- Warm transfer: the agent summarises the case to the human operator before connecting the customer.
- Cold transfer: the agent passes the call directly; the customer explains again.
Warm transfer reduces resolution time and improves CSAT. It requires integration with the switchboard via SIP trunk or the Twilio/3CX API. If your business uses AI agents for L1 support, this functionality is critical.
How to test transfer
- In the demo, ask for a case the agent cannot resolve (for example, a return outside policy).
- Check that the agent detects the limit and offers a transfer.
- Confirm that the call reaches the right queue and the operator receives context.
If the provider has no transfer or it is on the “roadmap”, the agent is an MVP, not a production solution.
Metrics and dashboard: measure what matters
Without metrics, you do not know whether the agent works. A useful dashboard shows:
- Resolution rate: % of calls closed without a human.
- CSAT: customer score after the call (1-5).
- P95 latency: 95th percentile of response time.
- Abandonment rate: % of customers who hang up before finishing.
- Average duration: minutes per call.
Typical benchmarks
- Resolution > 75%: the agent handles 3 out of every 4 calls.
- CSAT > 4/5: the customer is satisfied.
- P95 latency < 1.5 s: natural conversation.
- Abandonment < 10%: the flow does not frustrate.
If the provider does not expose these metrics on a dashboard, ask for them as CSV or via API. The lack of metrics is a sign that the system is not instrumented for production.
Frequently asked questions
What latency is acceptable in an AI voice agent?
Between 800 ms and 1.5 seconds from the customer finishing speaking to the agent responding. Below 800 ms the conversation feels natural; above 2 seconds the experience degrades and drop-offs increase. Latency depends on the full chain: transcription, LLM, voice synthesis and network.
Are synthetic or cloned voices better?
High-end synthetic voices (ElevenLabs Turbo v2.5, OpenAI Alloy) sound natural, handle emotions and cost less per call. Cloned voices of your team bring brand authenticity but require more audio samples, RGPD consent management and additional cost. Start with synthetic voices; clone only if voice identity is critical.
Can I integrate the voice agent with my CRM?
If the provider exposes a real-time webhook or REST API, yes. Well-designed agents send call events (start, end, transcript, outcome) to your CRM via HTTP POST or integrate directly with Salesforce, HubSpot or Pipedrive. Ask for the API documentation and an example payload before signing.
How much does an AI voice agent cost per month?
Typical SaaS providers charge between €0.10 and €0.40 per minute of call, plus a platform fee (€50-200/month). An SME with 200 calls averaging 3 minutes pays between €110 and €290/month. Custom solutions with their own infrastructure start from €800/month but reduce the variable cost per call.
Can the agent transfer calls to a human?
Good ones can. Look for warm transfer (the agent summarises the case before passing the call) or cold transfer (it transfers directly). Transfer requires integration with your switchboard (Asterisk, 3CX, Twilio) via SIP or API. Without this functionality, the agent is a dead end when it cannot resolve an issue.
How do you measure whether the agent is working well?
Three key metrics: resolution rate (percentage of calls closed without a human), customer CSAT score (post-call survey) and abandonment rate (customers who hang up before finishing). A well-configured agent exceeds 75% resolution, 4/5 CSAT and keeps abandonment below 10%. Demand a dashboard with these metrics before signing.
Next step
If you are comparing providers or already have a proposal on the table, these 7 criteria help you separate marketing demos from production-ready systems. The right decision depends on your volume, technology stack and use case.
STAKKER builds custom AI phone agents with latency < 1 s, native integrations with CRM and switchboard, and a real-time metrics dashboard. Each project starts with a free diagnosis where we audit your current call flow and design the architecture that fits.
If you want to assess whether an AI voice agent makes sense in your case, get in touch here.