Skip to content
Back to blog
10 min readBy The dialque Team

Seamless AI voice agent to human transfer with one continuous recording — the feature nobody else ships

dialqueAI now transfers live calls from AI to a human agent inside the same SIP session, with one continuous recording file for the entire conversation. No other AI voice provider does this — here's why it matters for QA, CSAT, AHT, and RBI FPC + IT Act evidentiary defensibility.

AI Voice AgentHuman TransferComplianceArchitecture

Every AI voice agent vendor sells you the same headline — "the AI can transfer to a human." What they don't put on the demo screen is what happens to the call session and the recording at the moment of transfer. In almost every stack shipping today, the AI hangs up, a brand new call is dialled to the human agent, and the recording splits into two files with a gap in the middle. That is a compliance defect, not a feature. From this week, dialqueAI ships the correct behaviour: one SIP session for the entire conversation, one continuous recording file, one transcript with speaker labels flipping from AI to AGENT mid-call. The customer never re-explains, the QA reviewer never stitches two files, and the RBI Fair Practices Code auditor gets exactly what the rule assumes — one call, one evidentiary artefact.

This post walks through what shipped, why every other AI voice platform gets this wrong, what it means for QA cost and legal defensibility, and how to switch it on inside your tenant.

What actually shipped

The feature is called seamless AI-to-human transfer, and it does five things end-to-end:

  • The AI voice agent takes the call (inbound or outbound), speaks in Hindi, English, or one of six regional Indian languages, and handles the conversation until either the customer explicitly asks for a human ("mujhe person se baat karni hai") or the LLM classifies the case as needing human judgement — compliance, negotiation, escalation, or a sensitive-topic trigger the tenant's prompt has flagged.
  • The transfer happens inside the same SIP session. There is a brief three-way bridge — AI, customer, human agent — during which the AI whispers a structured context handoff to the human ("collected: name Rakesh, loan account XXX9876, EMI overdue 2 months, wants restructure — over to you"), then drops off cleanly. No hangup, no re-dial, no hold music.
  • The human agent can be on a WebRTC browser softphone (our Agent Console) or on their mobile phone via the outbound trunk. Tenants pick the default; per-agent override is a single field — leave the mobile number blank and that agent falls back to browser softphone regardless of the tenant default.
  • Ring strategy is configurable per AI agent: round-robin (longest-idle first, hunts up to three agents in sequence) or ring-all (every available human rings, first to answer wins). Round-robin is the default because it keeps AHT distribution even across the human roster.
  • The recording is one continuous file — WAV or MP3 depending on tenant config — covering the AI turn and the human turn as a single audio artefact. The transcript is one file with speaker labels flipping mid-call. Billed seconds cover the full arc. The report row shows `TRANSFERRED_TO_HUMAN` as disposition with both leg durations in one line.

That's the whole feature in one paragraph. The interesting question is why nobody else in the AI voice category does it this way.

How dialqueAI compares to Bland, Retell, Vapi, Synthflow, ElevenLabs, PolyAI

This is the honest read. We have tested every one of these platforms in the last twelve months. What they call "human transfer" is not what an Indian BFSI compliance team means by the phrase.

| Platform | Transfer to human? | One SIP session? | Single recording file? | Context handoff to human? | |---|---|---|---|---| | Bland AI | Cold transfer (hangup + new call) | No — two calls | Two separate files | No | | Retell AI | Warm transfer via SIP REFER | Depends on trunk provider | Two files (per-participant recording) | Manual, via API | | Vapi | Supported via Twilio dial verb | No — Twilio spawns a child call | Two files, stitched by webhook | No native handoff | | Synthflow | Cold transfer only | No | Two files | No | | ElevenLabs Conversational AI | Beta — cold transfer | No | AI file ends, human file separate | No | | PolyAI | Warm transfer over their own SBC | Yes, on enterprise tier | Two files (participant-scoped recording) | Yes, structured | | dialqueAI | Warm transfer, in-session | Yes — single SIP session | One continuous file | Yes, whispered live |

The pattern is clear. Almost every AI voice provider outsources the SIP layer to a Twilio, Plivo, or Exotel wrapper — and those trunk APIs are per-call, not per-conversation. Once the AI wants to bring a human on, the wrapper spawns a new call leg with a new call-ID, and the recording engine — which is per-participant, per-call — writes a new file. The "AI provider" has no way to fix this because they don't own the SIP or recording path.

dialqueAI is built on a self-hosted Asterisk stack — we own the entire SIP + media path, the recording engine, and the RTP mixing bridge. That is what makes single-session, single-recording transfer physically possible. It isn't a feature you can bolt on to a hosted AI voice product; it's an architectural choice that has to be made before the first line of PBX config gets written.

How it works, step by step

Here's the actual sequence during a live transfer. This is what happens on the wire, not marketing simplification.

  1. Customer is talking to the AI on SIP session `A`. Recording is being written to `/recordings/<call-id>.wav` via Asterisk's `MixMonitor`, which captures the mixed audio of both legs of the bridge.
  2. The AI's LLM turn returns a `TRANSFER_TO_HUMAN` intent — either because the customer said so, or because the tenant's system prompt classified the utterance as needing escalation.
  3. The AI orchestrator opens a second leg to the first human in the ring queue — either via WebRTC (a WSS invite to their Agent Console browser tab) or via a SIP INVITE out to the outbound trunk if the human is on mobile.
  4. When the human answers, they land in a whisper channel first — the customer cannot hear this. The AI reads out the structured context summary in about 4-6 seconds ("Rakesh, loan XXX9876, EMI overdue 2 months, wants restructure"). The human confirms they're ready.
  5. The bridge briefly becomes three-party: AI + customer + human. The AI announces the handoff to the customer in the customer's own language ("main aapko humare senior agent Priya se connect kar raha hoon"), then unbridges itself.
  6. Session `A` is now a two-party bridge: customer + human. `MixMonitor` is still writing to the same file because it is scoped to the bridge, not to a participant.
  7. When either side hangs up, the recording finalises. One `.wav` (or `.mp3` if the tenant has post-processing enabled), one transcript, one report row, one billed duration covering the full arc.

The critical detail is step 6. In a hosted-trunk stack, when the AI leg detaches, the recording engine treats the AI's departure as end-of-call and finalises the file. The human leg is a new call, with a new recording. In our stack, the recording is bound to the bridge, not to the AI participant, so it survives the swap.

Business value: QA, CSAT, AHT, cost per contact

We ran this for four weeks internally and with two pilot tenants before shipping. The numbers that moved:

  • QA review time per call dropped 38%. Reviewers used to open two files, cue them up, and stitch mentally. Now they hit one play button and scrub across the AI-to-human boundary in the same waveform. On a 1,200-call weekly QA sample this is roughly 6-7 hours of reviewer time back.
  • CSAT held flat at the transfer boundary. In the two-file world, customers who got transferred showed a 14-point CSAT drop versus AI-only or human-only calls, driven almost entirely by "I had to explain everything again." With the context whisper, the human opens with "Hi Rakesh, I see your EMI is 2 months overdue and you want to talk about restructuring" — the drop collapses to statistical noise.
  • Average Handle Time on the human leg fell 22%. The human doesn't spend the first 90 seconds re-discovering what the AI already knew.
  • Cost per contact stays inside the tier price. dialque's per-second billing covers the full arc — AI seconds at your tier's AI rate, human seconds at the trunk rate — with no separate "transfer fee" or "orchestration fee" that hosted competitors typically charge because they're paying for two Twilio call legs instead of one. See [the pricing page](/pricing) for the tier-by-tier per-minute breakdown; Starter is ₹1,500, Growth ₹2,000, Enterprise ₹2,500 and the transfer feature is included on all three.

The AHT and CSAT numbers matter more than the QA time saving in most business cases, but the QA number is the one that shows up on a P&L immediately.

Compliance: RBI FPC, TRAI TCCCPR, IT Act evidentiary chain

This is the part that most AI voice vendors don't talk about because they can't fix it.

The RBI Fair Practices Code for lenders requires that all recovery communications be recorded and retained. The regulator's inspection guidance — and the standard for admissibility in DRT proceedings — assumes one call, one recording, one continuous artefact. When you hand a compliance officer two files with a gap between them, you have to prove the gap is genuine (session handoff) and not tampering. Opposing counsel in a borrower dispute will argue exactly the opposite. A single continuous file with a matching CDR row removes the argument.

The IT Act 2000 and the Information Technology (Reasonable Security Practices) Rules 2011 treat electronic records as evidence subject to Section 65B certification. Section 65B requires that the electronic record be produced from a computer that was operating properly at the time. Two files with independent hashes, produced by two different call sessions, each need their own 65B certificate. One file, one certificate — much cleaner in a litigation-hold scenario.

The TRAI TCCCPR 2018 rules on commercial communication don't directly speak to recording continuity, but the DND-and-consent audit trail is easier to defend when the full customer interaction is in one artefact — you can prove the consent capture happened on the AI leg and the substantive conversation happened on the human leg, all in one recording, with one timestamp continuum.

The DPDP Act 2023 data-minimisation posture also lands better with a single file: one retention policy, one deletion event, one access log entry per QA review. Two files means two of each, which is exactly what a data protection officer will flag on a privacy impact assessment.

None of this is theoretical. We have a Growth-tier NBFC tenant whose compliance team refused to onboard until the single-file behaviour was demonstrable. That refusal is what accelerated this feature to ship.

How to switch it on

The setup is in Super Admin, and it takes about ten minutes for a tenant that already has an AI agent live.

  1. Super Admin → Tenants → your tenant → Human Agents tab. Add each human with email, password, and (optionally) their mobile number.
  2. Set tenant default: WebRTC (Agent Console browser softphone) or Mobile (ring the phone via trunk).
  3. Per-agent override: if a specific agent should always use the browser regardless of the tenant default, leave their Phone Number blank.
  4. Edit AI Agent → check "Transfer to human" and pick a ring strategy (Round robin / Ring all).
  5. Add the transfer trigger to your system prompt: either rely on the explicit "I want a human" detection (on by default) or add your own escalation triggers ("If the customer mentions legal action, transfer immediately").

That's it. First transferred call shows up in the call log with `TRANSFERRED_TO_HUMAN` disposition, one recording, one transcript, one billed duration.

FAQ

Can I still use my own SIP trunk (BYOC) with human transfer?

Yes — the feature is transparent to the trunk layer. Whether you're on our included trunk, on TATA via our gateway, or on your own BYOC trunk (a Growth+ deployment choice), the SIP session continuity and single-recording behaviour work identically because it's an Asterisk-side property.

What happens if no human agent picks up?

The AI stays in the bridge. If ring-all times out or round-robin exhausts its three-hop hunt, the AI falls back to a configurable message — usually a callback promise, sometimes a voicemail-style capture — and closes the call. The recording still finalises as one file, disposition shows `TRANSFER_FAILED` for reporting.

Does the single recording work with real-time barge-in supervision?

Yes. A supervisor barging in on the human leg via Agent Console is added to the bridge as a silent participant; their audio isn't mixed into the recording unless the tenant explicitly enables supervisor-audio-in-recording (off by default for privacy reasons).

How does billing work across the AI and human legs?

Per-second billing continues across the entire session. AI seconds are billed at your tier's AI rate; human-leg seconds are billed at the human-agent rate (which includes the outbound trunk cost if the human is on mobile). One line on the invoice, one line in the report — see [the AI vs human cost analysis](/blog/ai-calling-bot-vs-human-agent-cost-conversion) for the arithmetic on when the mixed model beats human-only.

Is this available on the browser softphone AND mobile?

Both. The choice is a tenant default with per-agent override. See [the browser softphone vs mobile dialer comparison](/blog/browser-softphone-vs-mobile-phone-dialer) for the trade-offs — WebRTC gives you screen-pop and lower latency, mobile gives you agents-on-the-go and no data-plan dependency.

Can I test this without signing a contract?

Yes — every tier ships with a free trial that includes the transfer feature. Point an AI agent at a test scenario, hit transfer, download the single recording. That's the demo we recommend.

---

If you want to see the single-recording, single-session transfer running on your own use case, book a working session at [dialque.com/contact?source=demo&topic=seamless-ai-human-transfer](/contact?source=demo&topic=seamless-ai-human-transfer). We'll set up the AI agent, wire in a human on WebRTC or mobile, and hand you the recording file at the end so you can play it back to your compliance team.