{"count":17,"filters":{"task":null,"language":null,"streamingOnly":false},"_read_this_first":{"evidence":"Every perf claim carries an evidence field: measured | published | estimated. NOTHING is currently \"measured\". Do not quote estimated numbers to a customer.","licenseRisk":"clear = safe to ship. caution = read licenseNote. blocker = do not ship commercially until resolved.","streaming":"streaming:false means the model cannot be used in a live phone call, however good it sounds."},"languages_available":["as","bn","doi","gu","hi","ks","kok","mai","mr","ne","or","pa","sa","sd","ur","kn","ml","ta","te","brx","mni","sat","bho","mag","hne","en-IN"],"models":[{"id":"svara-tts-v1","task":"tts","displayName":"svara-TTS v1","author":"Kenpath Technologies","license":"Apache-2.0","licenseRisk":"caution","licenseNote":"Card states Apache-2.0. Architecture derives from Orpheus, which derives from a Llama backbone. Have counsel confirm the chain before a large contract.","params":"~3B","languages":["hi","bn","mr","te","kn","bho","mag","hne","mai","as","brx","doi","gu","ml","pa","ta","ne","sa","en-IN"],"streaming":true,"sampleRateHz":24000,"device":"gpu","vramGb":8,"perf":{"ttfbMs":793,"rtf":1.282,"evidence":"measured","note":"MEASURED on our A10G through vLLM: streaming time-to-first-audio 793ms (782-793ms across runs, variance of about 1ms), RTF 1.282. Before vLLM the same model on HuggingFace transformers took 2,473ms with RTF 2.099, so this is a 3.16x improvement and it is what makes live calls viable. Orpheus publishes ~180ms on an H100 and ~280ms on an A100 with vLLM, so 793ms on an A10G is consistent with the weaker GPU rather than a misconfiguration."},"voiceCloning":true,"emotionTags":["<happy>","<sad>","<anger>","<fear>","<clear>"],"caveats":["No Odia (or), Urdu (ur), Kashmiri (ks), Konkani (gom), Manipuri (mni), Santali (sat) or Sindhi (sd). Those seven can be transcribed and translated but not spoken.","ALL 38 VOICES VERIFIED: every voice was synthesised and quality-checked in the audit, 38/38 producing usable audio.","Emotion tags go at the END of the sentence, not the start.","Use <clear> for rupee amounts and account numbers — prioritises intelligibility over expression.","For HINDI, send Roman text (\"aapka order aa gaya\") rather than Devanagari. It sounds noticeably better. Other languages use native script.","Speaker IDs are per-language, e.g. \"Hindi (Female)\". Switching language changes the voice.","LoRA-friendly: this is the intended path for a custom brand voice.","24kHz output is downsampled to 8kHz mulaw for telephony by /v1/audio/speech.","MEASURED: cold start 147s from a cold boot, 32s warm. The API warms it automatically at startup so no request pays this.","MEASURED with vLLM: 6 concurrent syntheses in 13.9s against 83.7s serialised, a 6x improvement. Before vLLM capacity was roughly one.","The model occasionally runs away and generates overlong audio. A length bound plus a quality check catch it; X-Quality-Check reports the outcome and a degenerate stream falls back to a non-streaming retry.","The prompt format is easy to get wrong and fails silently — a wrong wrapper produces audio tokens that decode to near-silence rather than an error. Correct order is BOS, START_OF_HUMAN, AUDIO_TOKEN, text, END_OF_HUMAN, END_OF_TURN, START_OF_AI, START_OF_SPEECH.","vLLM must be given raw token_ids via prompt_token_ids and read back with detokenize=False. Decoding audio tokens with the TEXT tokenizer yields garbage — a well-known Orpheus trap."],"url":"https://huggingface.co/kenpath/svara-tts-v1","recommended":true},{"id":"indic-mio","task":"tts","displayName":"Indic-Mio","author":"SPRINGLab, IIT Madras","license":"UNLICENSED","licenseRisk":"blocker","licenseNote":"No licence stated on the model card. Base model (Aratako/MioTTS-0.6B) appears to be Apache-2.0. DO NOT ship commercially until SPRINGLab confirms in writing.","params":"0.6B","languages":["as","bn","brx","doi","gu","hi","kn","ks","kok","mai","ml","mni","mr","ne","or","pa","sa","sat","sd","ta","te","ur","en-IN"],"streaming":true,"sampleRateHz":44000,"device":"gpu","vramGb":2,"perf":{"rtf":0.1,"evidence":"published","note":"Card claims RTF < 0.1. Not reproduced on our A10G."},"voiceCloning":true,"emotionTags":["<happy>","<sad>","<angry>","<disgust>","<fear>","<surprise>","<enunciated>","<confused>","<whisper>"],"caveats":["LICENCE UNRESOLVED — this is the one blocker on an otherwise ideal model.","Only model covering all 22 scheduled languages plus English.","Zero-shot cloning via codec speaker embeddings keeps ONE voice consistent across every language. This is what makes mid-call language switching sound like the same human.","Recommended runtime is vLLM, which does not run well on Apple Silicon. Test this on the GPU box, not your Mac.","Stress a word with *asterisks*.","0.6B means roughly 5x more concurrent streams than svara-TTS for the same GPU."],"url":"https://huggingface.co/SPRINGLab/Indic-Mio"},{"id":"indic-parler-tts","task":"tts","displayName":"Indic Parler-TTS","author":"AI4Bharat","license":"Apache-2.0","licenseRisk":"clear","params":"~880M","languages":["as","bn","brx","doi","gu","hi","kn","kok","mai","ml","mni","mr","ne","or","sa","sat","sd","ta","te","ur","en-IN"],"streaming":false,"sampleRateHz":44100,"device":"gpu","vramGb":3,"perf":{"evidence":"estimated","note":"Autoregressive. Streaming is possible but nobody has wired it up. Treat as offline for now."},"voiceCloning":false,"caveats":["THIS IS WHY WE KEEP IT: it has Odia and Urdu, which svara-TTS lacks.","No Punjabi (pa) — svara covers that. The two models are complements.","Voices are selected by a natural-language description, not a speaker ID or reference clip.","Gated repo on Hugging Face — request access before you need it.","69 named voices available."],"url":"https://huggingface.co/ai4bharat/indic-parler-tts"},{"id":"supertonic-3","task":"tts","displayName":"Supertonic 3","author":"Supertone Inc.","license":"OpenRAIL-M","licenseRisk":"caution","licenseNote":"Weights are OpenRAIL-M: commercial use permitted, but with use-based restrictions (no impersonation without consent) and an attribution requirement. Sample code is MIT. Not equivalent to Apache-2.0.","params":"99M","languages":["hi","en-IN"],"streaming":true,"sampleRateHz":44100,"device":"cpu","vramGb":0,"perf":{"rtf":0.2,"evidence":"published","note":"RTF 0.200 on a 16-thread CPU (published, N=30). Our box has 8 vCPU so expect roughly 0.4 — still ~2.5x real time."},"voiceCloning":false,"caveats":["RUNS ON CPU. If TTS leaves the GPU, the GPU only serves STT and the LLM — roughly doubling concurrent calls. Architecturally the biggest single win available.","Hindi is the ONLY Indian language. No Tamil, Telugu, Bengali, Marathi.","Open weights ship fixed preset voices. Zero-shot cloning requires the vendor's paid Voice Builder.","Published benchmark beat Chatterbox Multilingual 4x on speed while Chatterbox ran on an RTX 3090."],"url":"https://huggingface.co/Supertone/supertonic-3"},{"id":"kokoro-82m","task":"tts","displayName":"Kokoro 82M","author":"hexgrad","license":"Apache-2.0","licenseRisk":"clear","params":"82M","languages":["hi","en-IN"],"streaming":true,"sampleRateHz":24000,"device":"either","vramGb":3,"perf":{"ttfbMs":300,"evidence":"published","note":"Project reports ~300ms first audio on GPU, 35-100x real time, ~3.1GB VRAM."},"voiceCloning":false,"caveats":["Very few Hindi voices. Fine as a fast fallback, not as your primary voice.","Scored 4.33/5 for Hindi naturalness in a third-party listening test (Aug 2026)."],"url":"https://huggingface.co/hexgrad/Kokoro-82M"},{"id":"indicf5","task":"tts","displayName":"IndicF5","author":"AI4Bharat","license":"UNLICENSED","licenseRisk":"blocker","licenseNote":"No licence tag on the model card. Only a terms-of-use note prohibiting non-consensual voice cloning. Safe for internal data generation; do not put it in a customer contract.","params":"~0.3B","languages":["as","bn","gu","hi","kn","ml","mr","or","pa","ta","te"],"streaming":false,"sampleRateHz":24000,"device":"gpu","vramGb":2,"perf":{"evidence":"published","note":"Flow-matching (F5). Generates the whole utterance before emitting anything. Unusable for live calls."},"voiceCloning":true,"caveats":["NOT FOR LIVE CALLS. No streaming — first audio equals full synthesis time.","NO ENGLISH. Useless for Hinglish on its own.","Best-in-class quality: scored 5.0/5 for Hindi naturalness in a third-party listening test, the only model to score perfect.","USE IT FOR: generating synthetic training data, pre-recorded prompts, and WhatsApp voice notes where latency does not matter.","Requires a reference clip AND the transcript of that clip."],"url":"https://huggingface.co/ai4bharat/IndicF5"},{"id":"indicconformer-600m","task":"stt","displayName":"IndicConformer 600M Multilingual","author":"AI4Bharat","license":"MIT","licenseRisk":"clear","params":"600M","languages":["as","bn","brx","doi","gu","hi","kn","ks","kok","mai","ml","mni","mr","ne","or","pa","sa","sat","sd","ta","te","ur"],"streaming":false,"sampleRateHz":16000,"device":"gpu","vramGb":2,"perf":{"evidence":"estimated","note":"Never benchmarked by us — we could not download it. See caveats."},"caveats":["BLOCKED: this is a GATED Hugging Face repo. Downloading returns 401 GatedRepoError. It needs an HF_TOKEN from an account that has accepted the licence on the model page. Approval is automatic (gated=auto), so it is one click — but somebody has to click it.","Every AI4Bharat and ARTPARK speech model we tried is gated the same way: IndicConformer, SraVaani, IndicWav2Vec.","WANTS 16kHz. Telephony is 8kHz, so you upsample on the way in. Accuracy cost is still UNMEASURED.","Offline model. Live streaming needs a chunking wrapper that does not exist yet — a real build task, not config.","Language code is a REQUIRED argument. There is no auto-detect. Pair it with Vaani-LID.","Two decoders: CTC (faster) and RNNT (more accurate). Benchmark both once you have access."],"url":"https://huggingface.co/ai4bharat/indic-conformer-600m-multilingual"},{"id":"whisper-large-v3-turbo","task":"stt","displayName":"Whisper large-v3-turbo","author":"OpenAI","license":"MIT","licenseRisk":"clear","params":"809M","languages":["hi","bn","mr","te","ta","kn","ml","gu","pa","or","as","ur","ne","sa","sd","en-IN"],"streaming":false,"sampleRateHz":16000,"device":"gpu","vramGb":2,"perf":{"ttfbMs":305,"rtf":0.04,"evidence":"measured","note":"MEASURED on real human speech (FLEURS, CC-BY-4.0) through the CTranslate2 int8 backend: 305ms per utterance, RTF 0.04. Word error rate 0.204 at 16kHz and 0.252 after a round trip through 8kHz mulaw; character error rate 0.099 and 0.121. Per language at 16kHz: Hindi WER 0.104, Telugu 0.219, Tamil 0.277. This is the default and the only model here fast enough for a live call."},"caveats":["THIS IS WHAT ACTUALLY RUNS TODAY, because IndicConformer is gated. Fully open weights, MIT licence.","MEASURED: the 8kHz telephony penalty is about 5 points of WER (0.204 to 0.252). This used to be the biggest open question in these docs and it turned out smaller than feared.","Genuinely good at code-switching, which matters more for Hinglish than raw per-language accuracy.","Offline model — transcribes a complete utterance, so it does not give you streaming partial transcripts. The live-call websocket emits one final transcript per turn.","FOR HINDI OR TELUGU SPECIFICALLY the dedicated fine-tunes are about twice as accurate: whisper-hindi-large-v2 scores WER 0.047 on Hindi against 0.104 here. They are roughly 15x slower (4.9s and 5.7s per utterance) and only decode their own language, so use them for transcription nobody is waiting on.","Run through CTranslate2 with int8_float16 rather than transformers: same weights, but the transformers backend measured 1695ms and was notably worse on Telugu (WER 0.734 against 0.219)."],"url":"https://huggingface.co/openai/whisper-large-v3-turbo","recommended":true},{"id":"sravaani","task":"stt","displayName":"SraVaani","author":"ARTPARK, IISc Bangalore","license":"UNVERIFIED","licenseRisk":"caution","licenseNote":"Not yet checked. Same team as Vaani-LID, which is MIT.","languages":["as","bn","brx","doi","gu","hi","kn","ks","kok","mai","ml","mni","mr","ne","or","pa","sa","sat","sd","ta","te","ur"],"streaming":false,"device":"gpu","perf":{"evidence":"published","note":"Authors claim 63 Indian languages and dialects. Untested by us."},"caveats":["UNTESTED. Listed because it may simply be better than IndicConformer — same team produced the strongest Indic LID we found.","Benchmark it head-to-head before committing to IndicConformer."],"url":"https://huggingface.co/blog/ARTPARK-IISc/vaani-lid"},{"id":"vaani-lid-v0","task":"lid","displayName":"Vaani-LID v0","author":"ARTPARK, IISc Bangalore","license":"MIT","licenseRisk":"clear","languages":["as","bn","brx","doi","gu","hi","kn","ks","kok","mai","ml","mni","mr","ne","or","pa","sa","sat","sd","ta","te","ur","bho","mag","hne","en-IN"],"streaming":false,"sampleRateHz":16000,"device":"gpu","vramGb":2,"perf":{"ttfbMs":198,"evidence":"measured","note":"MEASURED on our A10G: 198ms median per classification, including HTTP round trip. Fast enough to run on every turn."},"caveats":["42 Indian languages across four families. No gating, MIT licensed — the one AI4Bharat-adjacent model that is NOT gated.","Trained on 16kHz. Accuracy on 8kHz phone audio is still UNMEASURED.","Telugu vs Tamil is an EASY pair (Dravidian, 85.9% published). Hindi vs Urdu is near-impossible.","CONFIRMED IN OUR OWN TESTING: fed clean synthesised Hindi, it ranked Urdu 0.368 ABOVE Hindi 0.304. The published 58.7% Central Indo-Aryan figure is real, not theoretical. Do not build on Hindi/Urdu separation.","Authors advise freezing the encoder when deployment audio differs from training audio — which is your case.","Run it in PARALLEL with the STT, not before it. Then it costs zero added latency.","The published one-line pipeline() loader does NOT work on transformers 5.x: the model calls WhisperEncoder.from_pretrained inside its own __init__, which the meta-device loader rejects. Our server loads it manually instead."],"url":"https://huggingface.co/ARTPARK-IISc/Vaani-LID_v0","recommended":true},{"id":"qwen3-8b","task":"llm","displayName":"Qwen3 8B","author":"Alibaba","license":"Apache-2.0","licenseRisk":"clear","params":"8B","languages":["hi","en-IN","bn","ta","te","mr","gu"],"streaming":true,"device":"gpu","vramGb":6,"perf":{"evidence":"estimated","note":"NOT DEPLOYED and not planned on this hardware. Never measured here."},"caveats":["ABANDONED FOR THIS DEPLOYMENT, and the reason is structural rather than a preference. Serving a local LLM would need a second vLLM instance, and two vLLM instances on one GPU is broken (vllm issue #10643: the second instance miscounts free memory and ends up with a negative KV cache size). The single A10G already gives ~12.6GB to vLLM for text-to-speech.","The LLM is served through Bedrock instead (nova-lite, ~387ms measured). Its ~400ms is hidden behind the filler cache, which costs 1ms, so the local-LLM win would have been small anyway.","Would become worth revisiting on a second GPU, where it removes a network hop and the per-token cost."],"url":"https://huggingface.co/Qwen/Qwen3-8B"},{"id":"sarvam-m","task":"llm","displayName":"Sarvam-M","author":"Sarvam AI","license":"Apache-2.0","licenseRisk":"clear","licenseNote":"Apache-2.0 inherited from Mistral-Small-3.1-24B base.","params":"24B","languages":["hi","en-IN","bn","ta","te","mr","gu","kn","ml","pa","or"],"streaming":true,"device":"gpu","vramGb":14,"perf":{"evidence":"estimated"},"caveats":["Handles Indic scripts AND romanised (Hinglish) input natively.","14GB at Q4 leaves little room for STT and TTS on a 24GB card. Tight."],"url":"https://huggingface.co/sarvamai/sarvam-m"},{"id":"bedrock-claude-haiku-4-5","task":"llm","displayName":"Claude Haiku 4.5 (Bedrock)","author":"Anthropic via AWS","license":"Commercial API","licenseRisk":"clear","languages":["hi","en-IN","bn","ta","te","mr","gu","kn","ml","pa","or"],"streaming":true,"device":"api","perf":{"evidence":"estimated","note":"Not measurable yet — see caveats."},"caveats":["BLOCKED on this AWS account: Bedrock returns ResourceNotFoundException with \"Model use case details have not been submitted for this account\". Somebody must fill in the Anthropic use-case form in the Bedrock console. One-time, then wait ~15 minutes.","Also needs the cross-region inference profile id (us.anthropic.claude-haiku-4-5-...), not the bare model id. The bare id returns a throughput error."],"url":"https://docs.aws.amazon.com/bedrock/"},{"id":"nova-lite","task":"llm","displayName":"Amazon Nova Lite (Bedrock)","author":"Amazon via AWS","license":"Commercial API","licenseRisk":"clear","languages":["hi","en-IN","bn","ta","te","mr","gu","kn","ml","pa","or"],"streaming":true,"device":"api","perf":{"ttfbMs":485,"evidence":"measured","note":"MEASURED end to end from our GPU box: 485ms median, 719ms p95, for a short Hinglish reply."},"caveats":["THIS IS THE LLM THAT WORKS TODAY. Verified on this AWS account with the bare model id, no extra forms.","Replies in reasonable Hinglish. Weaker than Claude at nuance, which matters more for a sales call than a support call.","Cheapest option per token of anything here."],"url":"https://docs.aws.amazon.com/bedrock/","recommended":true},{"id":"silero-vad","task":"vad","displayName":"Silero VAD","author":"Silero Team","license":"MIT","licenseRisk":"clear","languages":[],"streaming":true,"device":"cpu","perf":{"evidence":"published","note":"Single-digit milliseconds on CPU."},"caveats":["Silero VAD is MIT. Silero STT models are NON-COMMERCIAL. Do not confuse the two.","Not tuned for Indian acoustic conditions — background fans, traffic, and short acknowledgements like \"haan\"."],"url":"https://github.com/snakers4/silero-vad","recommended":true},{"id":"smart-turn","task":"turn","displayName":"Smart Turn v3","author":"Pipecat","license":"BSD-2-Clause","licenseRisk":"clear","languages":[],"streaming":true,"device":"cpu","perf":{"evidence":"estimated","note":"~50ms per turn."},"caveats":["Decides end-of-turn from grammar, tone and pace — not just silence. Closest thing to how a human knows you are finished.","Ships in Pipecat as LocalSmartTurnAnalyzerV3, enabled by default.","Pair with the asymmetric word gate so \"haan\" does not interrupt the bot."],"url":"https://github.com/pipecat-ai/smart-turn","recommended":true},{"id":"indictrans2","task":"translate","displayName":"IndicTrans2","author":"AI4Bharat","license":"MIT","licenseRisk":"clear","languages":["as","bn","brx","doi","gu","hi","kn","ks","kok","mai","ml","mni","mr","ne","or","pa","sa","sat","sd","ta","te","ur","en-IN"],"streaming":false,"device":"gpu","vramGb":3,"caveats":["All 22 scheduled languages, both directions, plus English."],"url":"https://github.com/AI4Bharat/IndicTrans2"}]}