Build the intelligence layer behind Cheerio AI’s enterprise automation: LLM pipelines, agentic workflows, RAG systems and the evaluation rigor that keeps them honest in production.
Tech
Full Time
Bangalore (HSR Layout)
Onsite
2+ years
As per previous CTC
About Cheerio AI
Cheerio AI is a profitable, VC-backed AI automation company building intelligent products for large enterprises. We are backed by Artha VC and by operators and founders from Unacademy, Reddit, Elevation Capital and Dr Vaidya’s, and we serve VFS Global, IndusInd Bank, Meta and Hero MotoCorp. Over the last two years our valuation has grown 5X, which makes ESOP ownership here a real path to financial upside rather than a line on an offer letter.
Role overview
You will build the intelligence layer that powers Cheerio AI’s enterprise automation products: LLM pipelines, agentic workflows, RAG systems, and the AI features that turn raw customer data into intelligent, personalised action at scale.
You will work directly with Hardeep, our engineering lead, designing systems where model selection, context architecture and evaluation rigor matter as much as code quality. This is not a role for someone who wraps an LLM in a loop. It is for someone who understands why that breaks in production and builds for what comes after.
Key responsibilities
Design and build multi-agent workflows using LLMs: planning agents, tool-calling pipelines and orchestration layers across WhatsApp, CRM and omnichannel campaigns.
Architect and maintain retrieval-augmented generation pipelines, vector databases and knowledge ingestion systems that give our AI products accurate, context-aware memory over enterprise data.
Build eval frameworks that measure accuracy, hallucination rate, latency and cost across LLM tasks, so we ship AI features that work at enterprise scale and not just in demos.
Own integrations with frontier model APIs, prompt versioning, context window optimisation and fallback strategies for production reliability.
Partner with product and engineering to scope what is achievable, surface tradeoffs, and ship features that solve real enterprise problems.
Identify when prompt engineering hits its ceiling and lead fine-tuning or domain adaptation for specialised tasks.
Instrument LLM systems with logging, tracing and monitoring so non-deterministic failures are debuggable.
Stay close to the frontier and translate relevant advances into product-applicable experiments on a tight feedback loop.
Our tech stack
AI and orchestration: LangGraph for multi-agent workflows, Anthropic Claude via API as the primary LLM, Claude Code and Antigravity for AI-assisted development.
Search and retrieval: AWS OpenSearch for full-text and semantic search, PGVector for vector similarity inside Postgres.
Backend and infrastructure: Next.js, AWS Lambda, Postgres, Supabase.
Design and prototyping: Lovable for design reference and rapid UI scaffolding.
What we are looking for
2+ years of engineering experience with hands-on work building LLM-powered features in production, not just prototypes.
Deep familiarity with LLM APIs, prompt engineering and context management: token budgets, system prompt construction, tool use, structured output extraction.
Hands-on RAG experience: chunking strategies, embedding models, vector search and retrieval quality evaluation.
Experience with agentic frameworks such as LangGraph, CrewAI, AutoGen or custom orchestration.
Strong Python skills and comfort with async, streaming and the operational concerns of high-throughput LLM applications.
Working knowledge of ML infrastructure: model serving, vector databases, cloud deployment and cost optimisation for inference-heavy workloads.
Experience designing and running evals as a first-class engineering practice.
Ability to communicate model tradeoffs across cost, latency and quality to engineers, PMs and non-technical stakeholders.