01
Parameter-efficient fine-tuning
LoRA and QLoRA training on domain vocabulary, house style and strict output formats, with a held-out test set to show it worked.
Senior / Staff
MLOps engineers who fine-tune open-weight models such as Llama, Mistral and DeepSeek, and serve them on your own GPUs when data residency or token volume makes commercial APIs the wrong fit.
Role profile
Seniority and experience are agreed in the proposal, and you interview every engineer before they start. These are the skills that interview should test.
01
LoRA and QLoRA training on domain vocabulary, house style and strict output formats, with a held-out test set to show it worked.
02
PagedAttention, continuous batching and tensor parallelism tuned to your latency and throughput targets.
03
Choosing GPU types and instance counts from measured throughput, and comparing the cost against the API bill it would replace.
Typical work
Examples of the scope this role is hired for. Your statement of work sets the actual deliverables and how they are accepted.
What this role works with day to day. Tell us your stack and we’ll say plainly which parts we can staff.
Related service
AI development servicesAI built for the compliance review — validated outputs, humans in the loop, results measured in cycle time.
How hiring works
01
The stack, the seniority you need, the hours you want covered, and any certifications the work requires. We reply within one business day.
02
A written proposal sets out who we'd put forward, their seniority and experience, the scope, and the terms. If we can't staff the role well, we say so.
03
Nobody starts on your codebase until you've interviewed them and agreed. Use the skills on the role page as your interview checklist.
04
Working-hours overlap is agreed for each engagement and written into the statement of work — the shared window, who shifts hours, and how handoffs work outside it.
QuantmHill provides development teams from India for Indian startups and SMEs. Role profiles describe project capabilities; availability is confirmed for your engagement. Our team works on India Standard Time. Project working hours, availability and handoff responsibilities are agreed before kickoff.
Need specific certifications? Tell us at the start and we'll confirm whether we can staff to that requirement before you sign. There are no recruiting fees.
FAQ
Answered the way we would on a call. If yours isn’t here, send it — we reply within one business day.
Use RAG for facts that change. Fine-tune to teach a model specialized vocabulary, a rigid output format or a house style. Many systems need both.
It depends on volume and utilization. GPUs cost the same whether they're busy or idle, so self-hosting pays off when traffic is steady and high. We model it from your actual token volumes before you commit to hardware.
Related roles
Tell us the stack, the seniority you need and the hours you want covered. We reply within one business day, and you interview every engineer before they start.