Indian startup fine-tunes NVIDIA’s Nemotron 3 Nano to power a multilingual voice recruiter, boosting speed, cutting costs and tailoring AI for blue-collar job seekers
Vahan.ai has worked with NVIDIA’s technical team to fine-tune the chipmaker’s Nemotron 3 Nano model for its voice-based AI recruiter, as the Indian startup looks to build a specialised AI system for blue-collar hiring.
The collaboration took place through NVIDIA’s Inception programme, which works with startups to help them use NVIDIA’s hardware, models, and software libraries. Vahan.ai was brought into the programme earlier this year and was encouraged to experiment with one of NVIDIA’s smaller models for its recruitment use case, Madhav Krishna, founder and CEO of Vahan.ai, told Fortune India.
Vahan.ai’s AI recruiter speaks to blue-collar workers over phone calls, helping them discover jobs, apply and schedule interviews. It currently operates in English and Hindi, while the company plans to add support for eight more Indian languages. Vahan says its platform has enabled more than 1.5 million job placements across 920-plus cities and connected nearly 50 million Indians with employment opportunities.
Why Vahan.ai moved to a smaller model
For the NVIDIA collaboration, Vahan fine-tuned Nemotron 3 Nano using its proprietary production data, including real, human-corrected recruitment conversations. The aim was to make the model better suited to recruitment-specific workflows, including conversations with candidates and function execution.
Krishna said Vahan had previously been using a 120-billion-parameter model. Nemotron 3 Nano, at around 30 billion parameters, is roughly one-fourth the size. “We were using a 120-billion-parameter model. The NVIDIA model that we have fine-tuned is a 30-billion-parameter model, so it is a lot smaller, roughly one-fourth the size,” Krishna said.
For Vahan, the smaller model was a better fit because its AI recruiter does not need to solve complex problems or perform advanced reasoning. It needs to understand what a candidate is saying, carry a conversation, answer questions, and guide the person through the recruitment process.
“In our case, we’re building an AI agent that can essentially converse with a blue-collar worker. There’s very little technical understanding that the model needs to have. It needs to be able to carry a conversation, ask questions and answer the candidate’s questions,” Krishna said.
The company says the fine-tuned model outperformed both the base Nemotron model and its earlier cloud-hosted 120-billion-parameter model across seven recruitment-specific benchmarks. These included response correctness, function calling, human-like interactions, language matching and tool-argument accuracy. It also performed better on conversations it had not seen during training, according to Vahan.
Fine-tuning delivers faster responses and lower costs
The gains also showed up in the model’s response time. Vahan said the optimised deployment delivered nearly 6.7x faster time-to-first response and more than 3x lower average end-to-end response latency, while substantially reducing inference costs.
Krishna said the results were unexpected given the relatively small amount of work involved.
“Honestly, all the results were very surprising. We saw 3x improvement in most of the important benchmarks, nearly 7x faster time-to-first response and around 3x reduction in overall latency. Cost, I think, will also come down by 60% to 70% or more,” he said. The fine-tuning itself took Vahan’s team around two to three weeks, according to Krishna.
What made the biggest difference was not simply moving to a smaller model, he said, but feeding it data specific to the company’s business. “Definitely the data, the fine-tuning,” Krishna said when asked what contributed most to the improvement. “We benchmarked the model before fine-tuning, and it wasn’t doing very well. After we fine-tuned it, the results improved significantly.”
That data comes from thousands of hours of interactions between recruiters and candidates accumulated through Vahan’s recruitment workflows. The company is now looking to use similar data as it expands into more Indian languages.
Vahan currently works with English and Hindi and is looking at languages including Tamil, Telugu, and Marathi. Krishna acknowledged that expanding into these languages will bring new challenges around dialects, context, and cultural nuances. But he said the company already has recruitment data in these languages that can be used for further fine-tuning.
The NVIDIA collaboration also forms part of a broader plan at Vahan to develop specialised Small Language Models for recruitment. The company has set up a continuous evaluation and fine-tuning pipeline that uses production conversations to improve its models over time.
Krishna believes this approach could have implications beyond Vahan. For India, he argues, smaller specialised models could make more sense for many enterprise applications than relying on increasingly large general-purpose models.
“For India to become a notable force in AI, I think this is the approach to take,” he said. “Our datasets remain ours. Companies and organisations that have data can leverage that data and keep ownership of it. It doesn’t have to go to a large AI company.”
There is also an economics argument. Smaller models require less compute, while India needs to deploy AI applications at a scale that can serve a large and diverse population. “Smaller models running in India on less compute are a lot more efficient. We’ll be spending a lot less as a country,” Krishna said.