Applied Scientist at Microsoft, 4 YoE. I ship LLM agents end-to-end — synthetic data → SFT/RLHF post-training → eval infra. What I've owned:
- Computer Use Agent — synthetic data, SFT + RLHF, AI-safety alignment; task pass rate 45% → 78% across two model iterations
- Data Analyst Agent — synthetic data + post-training for tool use, multi-step file analysis, grounded reasoning
- Outlook Agent — reward modeling + RLHF post-training; function calling, email triage, multi-turn actions
- LLM evaluation platform serving 500+ internal Microsoft teams (benchmarking, personalization, grounding evals)
- AgenticLeo agent evaluation framework — improved alignment with human judgment by 25%
---
Looking for Founding AI Engineer / Applied AI Engineer roles at Seed–Series A AI startups. Especially interested in: coding agents, agent reliability + eval infra, agentic dev tools, vertical agents (legal / health / ops).
Prior: Data Science Intern @ CRED (session modeling; personalization AUC 0.84 → 0.91), Research Intern @ Microsoft Research (RAG over NCERT textbooks + LoRA fine-tuning). NSUT ECE 2023.
Signals: Kaggle x3 Expert (313 comp / 601 datasets / 1210 notebooks), 4th of 3300+ in Amazon ML Challenge 2023, 6th of 3300+ in Amazon ML Challenge 2021, Leetcode 2062 (top 1.92%).
Remote: Yes — 5–6 hr US overlap
Willing to relocate: Yes — SF Bay Area, NYC, anywhere in the US
Technologies: Python, PyTorch, LLM post-training (SFT / RLHF / DPO / GRPO), agent evaluation, RAG, reward modeling, Hugging Face, vLLM, LangChain, vector DBs, C++, SQL
---
Applied Scientist at Microsoft, 4 YoE. I ship LLM agents end-to-end — synthetic data → SFT/RLHF post-training → eval infra. What I've owned:
- Computer Use Agent — synthetic data, SFT + RLHF, AI-safety alignment; task pass rate 45% → 78% across two model iterations
- Data Analyst Agent — synthetic data + post-training for tool use, multi-step file analysis, grounded reasoning
- Outlook Agent — reward modeling + RLHF post-training; function calling, email triage, multi-turn actions
- LLM evaluation platform serving 500+ internal Microsoft teams (benchmarking, personalization, grounding evals)
- AgenticLeo agent evaluation framework — improved alignment with human judgment by 25%
---
Looking for Founding AI Engineer / Applied AI Engineer roles at Seed–Series A AI startups. Especially interested in: coding agents, agent reliability + eval infra, agentic dev tools, vertical agents (legal / health / ops).
Prior: Data Science Intern @ CRED (session modeling; personalization AUC 0.84 → 0.91), Research Intern @ Microsoft Research (RAG over NCERT textbooks + LoRA fine-tuning). NSUT ECE 2023.
Signals: Kaggle x3 Expert (313 comp / 601 datasets / 1210 notebooks), 4th of 3300+ in Amazon ML Challenge 2023, 6th of 3300+ in Amazon ML Challenge 2021, Leetcode 2062 (top 1.92%).
---
Résumé: https://drive.google.com/file/d/1FVxVAA5DpPfy2VNO-_k56beF6v3...
GitHub: https://github.com/shivammittal274
LinkedIn: https://linkedin.com/in/shivam274
Kaggle: https://www.kaggle.com/shivammittal274
Email: mittal.shivam103@gmail.com