How to set up AI voice agent transfer to human — a practical guide
The AI-to-human transfer works only as well as the six or seven decisions you make before your first live call — escalation triggers, ring strategy, audio path, context whisper, recording continuity, and fallback. Here is the decision tree a contact-centre ops lead has to walk, then the five-step setup on dialque.
If you are enabling AI-to-human transfer on your voice stack, the transfer will only be as good as the six or seven decisions you make before your first live call. Vendors show you a demo where the AI hands off gracefully, the human picks up in one ring, and the customer never notices — that demo is a rehearsed happy path. In production the same feature fails on unclear escalation triggers, a ring strategy nobody thought through, a human agent who could not hear the customer clearly, a context handoff that dropped the account number, or a "no human available" moment nobody planned for. This guide is the decision tree a contact-centre ops lead has to walk before setup, then the concrete steps to configure it on [dialque](/ai-voice-agents).
Decision 1 — Which calls should the AI escalate, and when?
The default failure mode is "escalate too eagerly" or "never escalate." Both destroy the economics. Sit with your team and lock in the exact triggers. In an Indian collections, insurance-servicing, or BFSI presales book, the working list is usually:
- Explicit customer request — "mujhe insaan se baat karni hai", "let me speak to a manager", "transfer me to a person." Detected in Hindi, English, and any of the regional Indian languages the AI supports. Hard trigger, always on.
- LLM-detected complexity — the customer says something outside the AI's competence: "I want to negotiate my settlement", "I need a moratorium for six months", "I want to file a complaint against your recovery officer". Your system prompt should list these explicitly for your book.
- Verification failure — the AI cannot confirm the borrower's identity via non-identifying verifiers (date of birth, PIN code, last four of the account) after two attempts. Escalate rather than disclose account information — this is an RBI Fair Practices Code guardrail that regulators actively test.
- High-value account flag — a customer flagged in your CRM as premium, high-outstanding, or legal-hold. The AI handles the greeting and consent capture, then hands off before any substantive conversation.
- Sentiment escalation — anger, distress, self-harm-adjacent statements. Detect, disengage cleanly, hand to a trained human.
- Compliance-sensitive triggers — the customer mentions RBI, ombudsman, legal action, media, or a specific regulator. These calls have to reach a human trained on the escalation protocol.
Write these into the AI agent's system prompt as explicit rules, not vague guidance. "If the customer mentions any of the following, transfer immediately: legal action, RBI complaint, moratorium, restructure, ombudsman." Precision here is what makes the AI escalate the right calls and hold the rest.
Decision 2 — Ring strategy: round-robin or ring-all?
Only two ring strategies matter in practice. Pick one per AI agent — you can mix strategies across campaigns.
| Factor | Round-robin (longest-idle first, hunts up to 3) | Ring-all (every available agent rings, first-to-answer wins) | |---|---|---| | Distribution of load | Even across the roster | Uneven — fastest fingers wins the day | | Answer speed | 8–15 seconds typical | 2–5 seconds typical | | Small team (2–5 humans) | Works — hunt exhausts quickly | Fine, but overkill | | Larger team (15+ humans) | Load spreads cleanly | Chaotic — every headset lights up on every transfer | | Impact on wrap-up time | Predictable | Erratic; ring-all creates "was that mine?" confusion | | Best fit | Collections, service desks, most Indian BFSI floors | Small high-priority sales inbound, VIP support queues |
Round-robin is the safer default. It keeps average handle time predictable and roster fatigue balanced. Ring-all fits queues where the SLA on pickup is under five seconds and the team is small enough to absorb the interrupt.
Decision 3 — Human agent audio path: browser or mobile?
The human agent has to hear the customer clearly and take notes at the same time. There are two lanes; pick a tenant default and allow per-agent overrides.
- Browser softphone (WebRTC). The human takes the transfer inside the Agent Console tab, hears the customer through a USB headset, and sees the AI's live transcript, disposition history, and CRM screen-pop in the same window. Best for office teams, best for anyone taking more than ten calls a day, best when you need reliable disposition capture and QA replay.
- Mobile phone. The human's mobile rings, they pick up, the customer is on the line. Best for field officers on a bike, home-lending verification agents walking a suburb, insurance surveyors on a rooftop. Also best for agents on a laptop with 2 GB of RAM where a browser softphone stutters.
The per-agent override on dialque is a single field on the profile — the mobile phone number. Populate it to route the mobile lane; clear it to keep that agent on the browser regardless of tenant default. Deeper walkthrough at [/blog/browser-softphone-vs-mobile-phone-dialer](/blog/browser-softphone-vs-mobile-phone-dialer).
Decision 4 — Context handoff: what should the AI tell the human before they pick up?
The single biggest cause of customers repeating themselves after a transfer is missing context handoff. Insist on a structured, whispered summary that reaches the human in the two-second window before the customer hears them.
The whisper should carry, in about four to six seconds:
- Customer name and preferred language
- Account or loan identifier
- The reason the AI is transferring (explicit request, complexity trigger, compliance)
- Current sentiment (calm, frustrated, distressed)
- Any commitment captured — for collections, the promise-to-pay amount and date; for support, the ticket reference; for sales, the qualified intent
"Rakesh, Hindi speaker, loan account XXX9876, wants to discuss a six-month moratorium, calm tone, no PTP captured — over to you." The human opens with "Hi Rakesh, I understand you'd like to talk about a moratorium option" — the customer never re-explains, and the average handle time on the human leg drops by roughly a fifth. That single change is worth more than most feature tweaks a vendor will pitch you.
Decision 5 — Recording and transcript continuity
Before you sign anything, ask the vendor for a demo call recording of an AI-to-human transfer. Then ask one specific question: is that one file, or two?
Almost every AI voice platform in the market today produces two files with a gap between them, because the AI hands off by hanging up and starting a fresh call to the human. Bland, Retell, Vapi, Synthflow, and most ElevenLabs Conversational AI setups behave this way. The two-file pattern is fine for casual customer-service use cases; it fails for RBI FPC compliance, DRT-admissible collections evidence, IT Act Section 65B certification, and any QA workflow where the reviewer needs to scrub across the transfer boundary.
dialque produces one continuous recording file covering the AI portion and the human portion of the same conversation, with a single transcript carrying speaker labels that flip mid-call from AI to AGENT. Full explanation at [/blog/ai-voice-agent-human-transfer-single-recording-continuous-session](/blog/ai-voice-agent-human-transfer-single-recording-continuous-session).
Decision 6 — What happens when no human is available?
Plan the fallback before you go live. Four defensible options:
- Voicemail capture — the AI apologises, invites the customer to leave a short message, and drops it into the human queue as a follow-up task. Fine for lower-stakes support.
- Callback queue — the AI collects a preferred callback window ("can we call you back between 3 and 5 pm today?") and books the callback into the agent roster. Best for collections and premium support.
- WhatsApp fallback — the AI sends a DLT-compliant WhatsApp message with a booked callback link. Works well for younger, mobile-first segments.
- Next-day auto-callback — the AI captures the intent, schedules a callback for the next business day at the same time, and confirms verbally.
Never let the call fall silent. A dead-air moment after "let me transfer you to a manager" is the worst outcome for CSAT and the hardest thing to explain to a compliance reviewer six months later.
Setup on dialque — a five-step walkthrough
Configuration takes about ten minutes on a tenant that already has an AI agent live.
- Super Admin → Tenants → your tenant → Human Agents tab. Add each human with email, password, and (optionally) their mobile phone number in E.164 format (+91XXXXXXXXXX). Populate the mobile number for field agents; leave it blank for office agents who will stay on the browser softphone.
- System Settings → Telephony → default audio path. Set to Browser (WebRTC) or Mobile. This is your tenant baseline; the per-agent override on the profile lets you flip individuals without touching the default.
- AI Agent editor → Transfer to human toggle. Check it. Pick a ring strategy — Round robin (longest-idle first, hunts up to three) or Ring all. Round-robin is the safer default for most teams.
- Prompt edit — escalation triggers. In your AI agent's system prompt, add the exact triggers your team agreed on in Decision 1. Explicit-request detection is on by default; you add the domain-specific ones.
- Test with your own number. Point the AI at your own mobile, run a scripted scenario ("I want to speak to someone about my loan"), and confirm the transfer lands on a human, the whisper carries the right context, and the recording drops into your bucket as one file with one transcript.
That is the full configuration for a Starter, Growth, or Enterprise tenant — the transfer feature ships on all three tiers (Starter ₹1,500, Growth ₹2,000, Enterprise ₹2,500 per agent per month). See [/pricing](/pricing) for the full breakdown.
Testing checklist before you go live
Run this before you point a live customer at the flow:
- Explicit-request transfer works in Hindi, English, and each regional language you support
- LLM-triggered transfer works on at least three domain-specific triggers from your prompt
- Round-robin distributes across three consecutive test transfers to three different agents
- The whisper reaches the human before the customer bridge; the human hears it, the customer does not
- The recording is one continuous file covering both portions of the conversation
- The transcript flips speaker labels from AI to AGENT at the right moment
- Fallback path fires cleanly when every human is offline
- The report row shows one billed duration covering the full arc, not two
- CRM screen-pop lands on the human's console with the AI's captured fields already populated
Green on every row, then roll to a pilot campaign for a week before the full switchover. Deep-dive on the cost trade-off between AI and human legs at [/blog/ai-calling-bot-vs-human-agent-cost-conversion](/blog/ai-calling-bot-vs-human-agent-cost-conversion) — and on the collections-specific case at [/virtual-recovery-agent](/virtual-recovery-agent).
FAQ
How long does the whole setup take end-to-end?
About ten minutes on a tenant that already has an AI agent live. Roughly a day of decision work with your ops lead before that — Decisions 1 through 6 in this guide should be documented and signed off before anyone opens Super Admin.
Can we start with a small human team and scale up?
Yes. Two or three humans is enough to test the flow. Round-robin scales cleanly as you add agents; ring-all becomes noisy above about ten. There is no minimum roster size for the feature.
Do we need separate phone numbers for the human agents?
No. On the browser softphone lane, humans use the Agent Console without their own DID. On the mobile lane, the number in the profile is where the trunk rings — it can be a personal or work mobile, as long as it is DLT-clean and consent-scoped correctly for your book.
What happens if a human accepts the transfer and then loses their internet or their mobile signal drops?
The AI detects the drop and re-attempts the transfer according to the ring strategy — the next agent in the round-robin hunt, or another agent in the ring-all set. If the hunt exhausts, the Decision 6 fallback fires. The recording still finalises as one file; the report row records the drop and the re-attempt.
Is per-second billing continuous across the AI and human portions?
Yes. One billed duration covers the full arc, priced at your tier — Starter ₹1,500, Growth ₹2,000, Enterprise ₹2,500 per agent per month with per-second usage on top. See [/features](/features) for the full feature-to-tier mapping.
Can we run the transfer inside a DLT-compliant outbound campaign?
Yes. The transfer is orthogonal to DLT — your DLT header and consent posture apply to the outbound leg; the transfer to a human happens inside the same session and doesn't require a separate DLT registration.
---
Ready to configure AI-to-human transfer for your team? [Book a working session](/contact?source=demo&topic=how-to-set-up-ai-voice-agent-transfer-to-human-team) and we will walk your ops lead through each decision before you touch a setting.