arrow_back Back to freelances
F

AI Voice Call Cost Optimization

Freelancer

Share:
placeIN home_workRemote assignmentContract publicAggregated job · IN

eventPublished on Sep 10, 2026 · verifiedWe confirmed on Sep 11, 2026 that it's still live

Is this your business?

₹ 1.500 – ₹ 12.500 per project

About the job

We are developing a production-grade AI voice calling platform for the Indian market and are looking for an experienced Realtime Voice AI / Conversational AI / Telephony Engineer to optimize our existing architecture. Our main objective is: Target AI voice cost: approximately ₹3/minute per call while maintaining natural, human-like conversation and very low response latency. This is not just an API integration project. We need someone who understands the complete realtime voice pipeline and can benchmark, optimize and redesign the architecture where necessary. Technologies Already Tested We have already experimented with: OpenAI Realtime OpenAI voice models Sarvam AI STT/TTS ElevenLabs TTS Plivo Twilio SIP/telephony Streaming STT → LLM → TTS architectures Realtime speech-to-speech architecture Our current implementations work, but the cascade architecture introduces noticeable latency, while our current realtime implementation is more expensive than our target. We therefore need an expert to find the optimal architecture rather than simply replacing one API with another. Primary Goals 1. Reduce AI voice cost Current AI voice cost is higher than our target. We want to reach approximately: ₹3/minute AI runtime target while maintaining acceptable: Voice quality Conversation accuracy Response speed Indian language support Reliability Telephony cost, SIP/carrier charges, infrastructure and taxes can be calculated separately. 2. Extremely low perceived latency The agent should feel like a real human conversation. Target: Customer stops speaking AI starts responding as quickly as possible Ideally sub-1-second perceived response Streaming audio No long silence before AI responds Immediate interruption handling We specifically want to avoid architectures where: Speech → wait for complete STT → wait for LLM → wait for complete response → TTS → playback creates 2–4+ seconds of delay. Architecture We Want to Evaluate We are open to multiple architectures. Option A — Direct Realtime Speech-to-Speech Customer ↓ Indian Telephony / SIP ↓ Voice Gateway ↓ OpenAI Realtime ↓ Customer This is currently our preferred approach if we can achieve the required cost. We want the developer to benchmark the latest suitable OpenAI realtime model/configuration and determine whether it can achieve approximately ₹3/minute through optimization.

Keep reading for free

Create a free account to see the full job and apply.

  • badgePortfolio visible to companies
  • notificationsNew job alert by email
  • favoriteAlways free, no catch