Senior AI Engineer · IIT Madras · Hyderabad, India

AI systems that work after the demo.

I design, build and run AI for teams that need it to hold up in daily use: the model and its guardrails, the backend, the cloud, and the dashboard that shows what it costs. One workflow at a time, from the first conversation to a system your team relies on, for companies in India, the US, the UK, Europe and the Gulf.

  • LLM agents
  • RAG
  • Voice AI
  • Document AI
  • Fine-tuning
  • AWS & Azure
~70%less manual ticket work for a B2B support team, in production
600employees on an HR platform I designed, built and deployed
60recruiters work every day in an ATS I built, with AI résumé screening
800+engineers taught AI/ML at Masai School

Worked with

  • Bridgetown Consulting GroupBridgetown Consulting Group
  • Atlas SystemsAtlas Systems
  • Bluekyte.AI (Counsello AI)Bluekyte.AI (Counsello AI)
  • Open Data FabricOpen Data Fabric
  • Walker SandsWalker Sands
  • BalihansBalihans
  • ShaktyAIShaktyAI
  • MentionNowMentionNow
  • Masai SchoolMasai School
  • StatinferStatinfer
  • MedTourEasyMedTourEasy
  • VillageagroVillageagro
  • Mazagon Dock ShipbuildersMazagon Dock Shipbuilders
  • NDA before details
  • Built in your cloud account
  • Code and IP are yours
  • Evals and a cost dashboard in every build
  • Invoices in INR, USD, GBP, EUR or AED

Services · What I can build for you

Eight kinds of AI systems I've built

Each one solves a problem most industries share, and each one comes with the backend, the cloud and the monitoring around it. Pick your industry to see which ones fit your team's week.

Your industry

Email & ticket automation

Agents that read incoming messages, file them, catch duplicates and draft replies. ~70% less manual work in production.

Fits: support, IT services, logistics, insurance, SaaS

Document intelligence

Pull structured data out of PDFs, spreadsheets, scans and handwriting, across 10+ file formats.

Fits: finance, insurance, legal, healthcare admin, real estate

Knowledge assistants

Chat with your company's documents, organized by department, with answers grounded in the source files.

Fits: agencies, consulting, legal, any document-heavy team

Voice agents

Outbound and inbound calls that hold a natural conversation, cope with interruptions and log every answer.

Fits: healthcare, research, real estate, hospitality, collections

AI products with guardrails

Customer-facing assistants with strict output contracts, guardrails that escalate instead of guessing, failover when a model fails, and cost tracking per conversation.

Fits: healthtech, wellness, edtech, fintech apps

Risk & compliance review

Agent teams that run multi-stage reviews, with a checking agent that validates each stage before the next.

Fits: fintech, banks, SaaS vendors, procurement

Recruitment & HR automation

An ATS that runs from client and vendor onboarding through AI résumé screening to hiring, used by 60 people, and an HRMS for 600 employees with per-unit access control and AI attendance reports.

Fits: staffing and recruitment firms, HR teams, multi-entity groups

AI visibility

Measure and improve how ChatGPT, Claude and Perplexity describe your brand, with MentionNow, which I co-founded.

Fits: any brand, in any market

Selected work

What changed first, then what's under the hood

Five systems I built end to end. The outcome is up front for founders, the drawing and the decisions are there for engineers, and each card says how I know it works.

DWG 01B2B supportLiveAtlas Systems · 2024–26

~70% less manual ticket work for a B2B support team

The team read every email and call transcript by hand to open tickets under SLA. Now an AI workflow reads each message, files it, spots duplicates, links related tickets and drafts the reply. People review instead of retyping.

~70%less manual review and ticket creation
~40%less code per new integration
  • MCP tools over Outlook and SharePoint, reusable across integrations
  • Deduplication and parent/child links so automation doesn't flood the queue
  • Drafted replies stay with a person for review

How I knowThe ~70% is the reduction in manual review and ticket-creation effort the support team reported after rollout; the ~40% is the drop in integration code once the MCP tools were reused.

  • MCP
  • LangGraph
  • LangChain
  • Outlook
  • SharePoint

Full case study: next stepRole: AI/ML Engineer

Email · call transcript MCP tools: Outlook, SharePoint LangGraph workflow
CategorizeDeduplicateLink parent / child
Ticket + drafted reply

SHEET 01 · MESSAGE PATH

DWG 02HealthcareIn testing with doctorsBridgetown Consulting Group · 2026

A physiotherapy app that knows when to say "see a doctor"

PTBuddy helps people in the US manage physiotherapy between doctor visits. New users answer questions about their pain and habits, get a weekly exercise program built from a physiotherapist's video library, then log how the pain changes. The next week's program is re-ranked from what they report. Inside the app, an assistant called Anya already knows their history, so they can just ask. Its first rule is to escalate: anything beyond self-managed physio goes to a doctor, not a home remedy.

~1.5 stime to first token (p95)
~$0.005cost per conversation
~98%replies passing the JSON contract
  • Clinical guardrails: escalate to a doctor instead of guessing
  • A strict JSON response contract the app can trust
  • Rolling summaries keep memory inside a fixed token budget
  • Model failover that ends in a deterministic keyword fallback
  • Tokens, latency and cost logged for every call

How I knowEvery reply is checked against the JSON contract before the app renders it, and tokens, latency and cost are logged per call in CloudWatch, so quality and spend are visible for each conversation.

  • LangGraph
  • Gemini 2.5 Flash
  • SSE streaming
  • AWS
  • PostgreSQL
  • Stripe

Full case study: next stepRole: AI, backend and infrastructure

Onboarding questionnaire Weekly exercise program
Exercise videosAnya assistantEscalate to doctor
Weekly pain check-in Next program re-ranked

SHEET 02 · WEEKLY LOOP

DWG 03RecruitmentLiveBridgetown Consulting Group · 2026

One system that carries a recruitment firm from client to hire

The recruitment teams worked across separate tools from the first client conversation to a candidate's first day. Now one system runs the whole pipeline: clients and vendors are onboarded, their requirements become open jobs, candidates are onboarded and their documents collected, AI screens résumés against each requirement, applications are submitted and tracked, interviews are scheduled, and the hired candidate is onboarded. Sixty people across the teams work in the same place instead of chasing each other for status.

200candidates tracked
60people across the teams work in it
7stages, one system
  • AI résumé screening ranks candidates against each job's requirements
  • Client and vendor onboarding feeds requirements straight into the pipeline
  • Document collection and interview scheduling live in the same flow

How I knowIn daily use by 60 people across the recruitment teams, with 200 candidates tracked through all seven stages in the system itself.

  • FastAPI
  • PostgreSQL
  • AWS
  • GPT-4o

Full case study: next stepRole: design, database, backend, deployment

Client & vendor onboarding Job requirements Candidate onboarding · documents AI résumé screening Application · interview scheduling Candidate onboarded

SHEET 03 · HIRING PATH

DWG 04Voice AIDeliveredAtlas Systems · 2024–26

A survey caller that sounds like a conversation, not a phone menu

An outbound caller that asks survey questions in natural conversation, checks each answer, and copes when people interrupt or talk over it. The hard part was live-call behavior, so I tuned turn detection, interruption handling and transcription for real callers.

~60%of answered calls finish the survey
~800 msmedian response time
  • Twilio streams call audio over WebSockets to OpenAI Realtime
  • Prompt-driven questions, with every answer validated
  • Turn detection, voice and language tuned for live calls

How I knowEvery answer is validated before the next question is asked, and turn detection, interruption handling and transcription were tuned on real calls rather than test audio.

  • Twilio
  • OpenAI Realtime API
  • WebSockets
  • Python

Full case study: next stepRole: AI/ML Engineer

Twilio outbound call WebSocket audio stream OpenAI Realtime
Turn detectionInterruptionsAnswer check
Next question · survey record

SHEET 04 · CALL PATH

DWG 05LegalDeliveredBluekyte.AI (Counsello AI) · 2024

Teaching LLaMA 3 Indian law: 9.75M tokens, one A100

General models stumble on Indian legal language. I adapted LLaMA 3 8B with LoRA continual pre-training on 15 Indian law textbooks, from cleaning the text to a domain-adapted model.

9.75Mtraining tokens
15Indian law textbooks
20+ htraining on one A100
  • Continual pre-training with LoRA instead of full fine-tuning
  • Corpus cleaned and tokenized in-house
  • LLaMA 3 8B
  • LoRA
  • Hugging Face
  • NVIDIA A100

Full case study: next stepRole: AI Engineer

15 Indian law textbooks Clean + tokenize → 9.75M tokens LoRA continual pre-training · A100 LLaMA 3 8B, legal domain Eval: base vs tuned

SHEET 05 · TRAINING PATH

Also built

  • HR techMulti-entity HRMS600 employees across five business units: attendance, leave and administration, with role-based access, strict data isolation, AI-assisted leave and AI attendance reports.
  • Risk & complianceMulti-agent vendor risk reviewA six-stage review run by a CrewAI crew, with a QA agent checking each stage.
  • EnterpriseMultimodal RAG platform10+ file formats, agent routing and QA evaluation, deployed on Azure Container Apps.
  • Marketing agencyDepartment knowledge platformGoogle Drive RAG per department for Walker Sands, delivered through Balihans.
  • OperationsDocument intelligence pipelineOCR, Azure Document Intelligence and image understanding for documents and handwriting.
  • InfrastructureQwen 2.5 7B in productionAutoscaled inference on SageMaker and EC2.
  • Customer engagementKollect AISMS, WhatsApp, email and IVR on AWS serverless, ~30% more responsive after load testing.

How I build

Watch a production agent handle a bad day

A hull is judged in a storm, not on the slipway. Run the same assistant through four situations and see which part of the system does the work. The R-tags point to the six rules below.

Run scenario
R1Requestuser message
R3Guardrailsscope check
Escalatesee a doctor
R4Contextrolling summary
R2Routerpick a model
PrimaryGemini 2.5 Flash
Secondaryfailover model
R3Keyword fallbackdeterministic
R3JSON contractschema check
R6Response+ telemetry
First token1.46 s
Tokens1,184
Cost$0.0049
OutcomeAnswered

    Modelled on the PTBuddy assistant's architecture (DWG 02). Replay any scenario to see which part of the system does the work.

    1. Start with the business problemA measurable result, not a technology demo.
    2. Use the simplest thing that worksOne model call, RAG, a workflow, an agent or several, chosen on purpose.
    3. Build for failureStructured outputs, validation, retries, failover and fallbacks.
    4. Control context and costSummaries, bounded memory, token tracking and the right model for each step.
    5. Engineer the plumbingAPIs, queues, storage, CI/CD, logs and monitoring.
    6. Measure what changedTime saved, speed, reliability and cost, before and after.

    Engagements · Ways to work together

    Start small. Build once we both know it will work.

    Most engagements start with a short, fixed-scope Sprint or Audit. What it finds decides whether there is anything worth building, and we scope the build from there.

    Every build ships with
    • Structured outputs and validation
    • An eval set with pass/fail checks
    • Failover and fallbacks
    • Token and cost tracking per call
    • Logs, dashboards and alerts
    • Docs and a recorded handover

    2 weeks · fixed

    AI Opportunity Sprint

    For teams who know AI could help but not where to start.

    • Map one workflow and what it costs today
    • Prototype the riskiest part
    • Cost per task or conversation
    • A go/no-go and a build plan
    Discuss a Sprint

    6 weeks · two-week loops

    Agent Pilot

    For one workflow, taken all the way to a monitored pilot.

    • One agent taken to a monitored pilot
    • Working software every two weeks
    • Evals and a cost dashboard included
    • Runs in your cloud; full handover
    Plan a Pilot

    Monthly

    Fractional AI Engineer

    For founders who need a senior AI engineer, about a day a week.

    • Roadmap and architecture reviews
    • Eval and cost monitoring
    • Help hiring your AI team
    Ask about availability

    Free: an AI Visibility Snapshot for your brand

    A one-page report showing how ChatGPT, Claude and Perplexity describe your company next to your competitors. No call needed.

    Request a snapshot

    Terms · How I work

    Clear steps, written terms, and your data stays yours

    1. A 30-minute callWe talk through the workflow and what "working" means for you. You get a written proposal within two business days.
    2. A Sprint or an AuditTwo weeks, fixed price. You get a prototype or a ranked list of fixes, a cost model, and a clear go/no-go.
    3. Build in two-week loopsWorking software and a demo every two weeks, with evals and a cost dashboard. Full handover at the end.

    Communication

    • Replies within one business day
    • A written update every week
    • Your tools: Slack, Teams, email, or WhatsApp

    Contracts & IP

    • NDA before you share details
    • MSA and a statement of work, or your own paper
    • Code and IP are yours once paid

    Your data

    • Built in your AWS or Azure account and region
    • Model providers listed; no training on your data
    • Access removed and data deleted at the end

    Compliance-aware

    • DPA with SCCs (EU), UK IDTA, or a HIPAA BAA where needed
    • AI disclosure built into chatbots and voice agents
    • Your security questionnaire, answered honestly

    Billing

    • Invoices in INR, USD, GBP, EUR or AED
    • 50% to start a Sprint or Audit; milestones for builds
    • Bank transfer or Wise

    No lock-in

    • You get the code, docs and dashboards
    • A recorded walkthrough for your team
    • Optional monthly support, never required

    Questions clients ask

    Straight answers before the first call

    Do you work with companies outside India?
    Yes. I work remotely with teams in India, the US, the UK, Europe and the Gulf. Indian and Gulf teams share my whole working day, Europe most of the afternoon, and US calls fit my evening. Invoices go out in INR, USD, GBP, EUR or AED.
    Where does the system run, and who owns the code?
    In your own AWS or Azure account and region. The code, documentation and dashboards are yours once paid, with a recorded walkthrough for your team, and my access is removed at the end.
    Will my data be used to train AI models?
    No. Every model provider is listed up front and none of them trains on your data. Where it is needed I sign a DPA with SCCs for the EU, a UK IDTA or a HIPAA BAA, and an NDA before you share any details.
    How does an engagement start?
    With a 30-minute call and a written proposal within two business days. Most teams then run a two-week, fixed-price Sprint or Audit, which ends in a prototype or a ranked list of fixes, a cost model and a clear go/no-go. Builds run in two-week loops with working software each time.
    Which models and tools do you work with?
    OpenAI, Claude and Gemini models through their APIs, and open models such as LLaMA and Qwen where fine-tuning or self-hosting makes sense. LangGraph, LangChain, CrewAI and the Model Context Protocol for agents; Python, FastAPI and PostgreSQL for the backend; AWS and Azure for the cloud. Each step gets the simplest thing that works.
    Can you fix an AI system that already exists?
    Yes. The Production-Readiness Audit is for AI that is flaky, slow or expensive: I pull failure patterns from real conversations, set up an eval baseline with pass/fail checks, review failover and guardrails, and rank the cost and latency fixes.
    Are you available part-time or as a fractional AI engineer?
    Yes. As a fractional AI engineer I give a founder about a day a week for roadmap and architecture reviews, eval and cost monitoring, and help hiring an AI team.

    Working hours

    Where your working day meets mine

    I'm in Hyderabad (IST, UTC+5:30). Indian and Gulf teams share my whole day, Europe most of the afternoon, and US calls fit my evening.

    IndiaHyderabad · Bengaluru · Mumbai
    09:00–17:00 IST
    UAEDubai · Mon–Fri
    10:30–18:30 IST
    Saudi ArabiaRiyadh · Sun–Thu
    11:30–19:30 IST
    EuropeBerlin · Amsterdam
    12:30–20:30 IST
    United KingdomLondon
    13:30–21:30 IST
    US EastNew York
    18:30–02:30 IST
    US WestSan Francisco
    21:30–05:30 IST
    SingaporeUTC+8
    06:30–14:30 IST

    Your 9:00–17:00, in ISTMy core hoursNow in Hyderabad: Times as of early October 2026; US and European clocks change in late October and early November.

    About · Srinivas Dharavath

    From ship hulls to AI systems

    Srinivas Dharavath is a Senior AI Engineer and AI consultant in Hyderabad, India, specialising in production systems built on large language models (LLMs). He works remotely with startups, agencies and enterprises in India, the US, the UK, Europe and the Gulf. He designs, builds and runs production generative AI systems: AI agents and multi-agent workflows with LangGraph, CrewAI and the Model Context Protocol (MCP); retrieval-augmented generation (RAG) over company documents; real-time voice AI on Twilio and the OpenAI Realtime API; document intelligence with OCR and Azure Document Intelligence; LLM fine-tuning (LLaMA 3 with LoRA) and LLM evaluation; and recruitment and HR automation, including an AI-enabled ATS and a multi-entity HRMS. His systems run on AWS and Azure with Python, FastAPI and PostgreSQL. He holds a B.Tech and M.Tech from IIT Madras, teaches AI/ML engineering at Masai School, and co-founded MentionNow, an AI visibility (GEO) platform. He takes fixed-scope consulting engagements and fractional AI engineering roles.

    I trained as a naval architect at IIT Madras, where you design hulls to stay stable under loads you can't fully predict. Now I do the same for AI systems.

    Over four years I've gone from AWS backends to fine-tuning LLaMA 3 on Indian law, and then to shipping agents, RAG and real-time voice AI for healthcare, recruitment, risk & compliance and enterprise teams. I own the whole path, from requirements and architecture to deployment and monitoring, and I track what matters after launch: latency, cost and failure rate.

    I also teach AI/ML engineering to 800+ students at Masai School, and I co-founded MentionNow, which measures how AI engines talk about brands.

    Outside work I make short films. The planning, the shot list and the edit turn out to need the same discipline as building a system.

    1. Senior AI Engineer · Bridgetown Consulting GroupPTBuddy physiotherapy app with the Anya assistant, a multi-entity HRMS and an AI-enabled ATS
    2. AI/ML Engineer · Atlas SystemsRisk-review agents, voice AI, enterprise RAG, MCP automation
    3. AI Engineer · Bluekyte.AI (Counsello AI)LLaMA 3 legal fine-tune, document AI, Qwen deployment
    4. Application Developer & AWS Solution Architect · Open Data FabricServerless engagement platform, APIs, infrastructure as code
    5. B.Tech + M.Tech · IIT MadrasNaval Architecture & Ocean Engineering

    Beyond the day job

    • Co-founder, MentionNow: AI visibility for brands, with clients in the US and Malaysia
    • AI/ML Teaching Assistant, Masai School: 800+ students
    • AI Engineer (Consultant), ShaktyAI: RAG agents and multi-agent systems
    • AI Engineer (Consultant), Walker Sands: enterprise agentic systems and workflow automation

    Tools I ship with

    Agents & models

    • LangGraph
    • LangChain
    • CrewAI
    • MCP
    • OpenAI
    • Claude
    • Gemini
    • Llama
    • Hugging Face

    Backend & data

    • Python
    • FastAPI
    • Django
    • PostgreSQL
    • MongoDB
    • Supabase
    • ChromaDB

    Cloud & ops

    • AWS
    • Azure
    • Docker
    • Kubernetes
    • Terraform

    Channels & product

    • Twilio
    • WhatsApp
    • Stripe
    • React
    • Next.js

    Get in touch

    Tell me which workflow should work better.

    Pick a 30-minute slot on my calendar, or send a short brief. I reply within one business day with questions or a rough plan. Happy to sign an NDA first.

    Your details are used only to reply to you.