AI Platform
/Senior leadership
Staff Infra Engineer — LLM Serving
Inference & Serving Systems
Compensation
INR 45 - 75 LPA
Engagement
Full-Time
Permanent role. Full-time commitment. This is an on-site role at Kolkata. It is not remote and not hybrid.
Scope of role
Set direction within your domain. Build and mentor a team. Own outcomes at the function level.
01 — The role
Why this role exists at EduRankAI
Own the inference fleet. Sub-100ms first-token latency at our scale, on infrastructure we control. Quantisation, batching, KV cache, scheduling, autoscaling — everything between "model checkpoint" and "API response". You will choose the inference stack, defend the choice, and own the cost-per-token reporting that finance uses to plan. When something goes wrong at 2am, you are who the on-call calls.
02 — The work
What you will own
- 01 Own the inference stack end-to-end (vLLM / SGLang / TensorRT-LLM).
- 02 Own the quantisation pipeline (GPTQ / AWQ / FP8) and the regression eval that says it shipped safely.
- 03 Continuous batching and KV-cache management — including the unpopular hard cases.
- 04 Multi-region serving and autoscaling.
- 05 Cost-per-token observability used by finance.
- 06 SLO ownership for every AI feature that ships.
03 — The expertise
What we look for
04 — The bar
Who thrives here
- → You have operated an inference fleet that served >1000 RPS in production for at least 12 months.
- → You can describe a specific quantisation regression you caught and how.
- → You have written at least one CUDA or Triton kernel that shipped to production.
- → You read kernel logs the way other people read newspapers.
- → You can defend a one-pager that argues for owning the inference stack instead of buying it.
Terms of Engagement
How working time works at this level
No counted hours. A five-day week with a four-hour overlap window, plus responsibility for the coverage of your area — including who is on call and when.
Type
Lead
Working days
5 per week
Rest days
2 per week
Hours counted
No
Daily overlap
4h
Measured by
The outcomes of the area you lead, the reliability of its coverage, and the growth of the people in it.
Where this stands legally
Hours are not counted at this level. The statutory ceiling of 9 hours a day and 48 a week remains the limit the role is designed to fit inside: if the work cannot be done within a normal week, the scope is wrong, not the person.
The full per-level model is published at Working Hours by Level.
05 — Hiring process
What to expect after you apply
- 01
Application review
Every application is read personally within five business days. We respond either way.
- 02
Take-home or live exercise
Role-specific. Time-boxed. Real problems we are actually working on, not invented puzzles.
- 03
Conversations
Deep technical and values conversations with the team you would join. No trick questions. No panel ambushes.
- 04
Offer or honest no
If yes: digital offer letter, signed in-portal, transparent terms. If no: written feedback if you want it.
Before you start
What we will collect. What it costs. What we will not do with it.
We will collect
- Name, email, phone — Account + application updates. No marketing.
- Resume / portfolio link — Human review of your work.
- Date + place of birth — Identity verification only.
- Your written responses — Selection rubric. Read by humans.
- Government ID (later) — Anti-fraud at offer / interview stage. Not at signup.
We will never
- Sell your data
- Share with third-party recruiters
- Use for advertising
- Train models on it
- Send marketing email
Our situation
EduRankAI is a small, independent organization building long-term capabilities in educational intelligence, advanced AI systems, and research infrastructure. We take no advertiser money, no donations with strings attached, and no investor pressure on hiring decisions. Applying is free, and every application is read by a human — recruitment, technical, academic and leadership teams. It buys us the right to be honest.
Ready to apply?
We read every application personally. If you are the right person for this role — regardless of pedigree, background, or where you are based — you will hear back from us within five business days.
Other roles in AI Platform
Explore related openings
Lead Full-Time