Available Now

Member of Technical Staff, Applied AI

GC AI is building the legal AI platform that 1,800+ companies rely on. I have been building the same multi-model LLM orchestration, RAG pipelines, and eval-gated production AI for 7+ years. FacadeDriver routes 30+ foundation models. LLM-as-Judge gates block production. 1,690 ground-truth samples.

95% Match20267+ yearsHouston, TX - Remote
Sai Likhith Kanuparthi

Relevant Work Experience Explore All ›

Open Source Explore All ›

// 02 - About

I'm the engineer your ML platform team calls when the foundation model demo worked once and now production is on fire.

My day-to-day at Airbnb is the unglamorous half of AI: multi-model isolation, routing across 30+ LLMs, retry and fallback when a provider rate-limits you at 2 AM, PII detection that doesn't false-positive on the word "John," and evaluation frameworks that catch drift before users do.

I've spent 7+ years migrating between stacks - Java at Oracle, PySpark at Shell, Kafka at Southwest, Spring Boot & Kubernetes at Eli Lilly, and now Python / FastAPI / Airflow / AWS Bedrock at Airbnb. What persists is a taste for systems that fail gracefully, observe themselves, and don't wake humans at 3 AM.

GC AI is building the legal AI platform that 1,800+ companies rely on. They need multi-model LLM orchestration, legal-grade evals, RAG pipelines for complex document analysis, and rapid prototyping frameworks. That is exactly what I build at Airbnb today. I want to bring this to GC AI.

- Sai

Multi-tenant ML platform

Built Airbnb's BPI Virtual Analyst - 4 partner teams, 128+ concurrent internal devs, 30+ LLMs behind one interface. Provider lock-in is a config flag.

16x scale on workflows

Refactored the analysis pipeline from 600 rows to 10,000 per run. Same hardware, batching & chunked uploads.

23-version eval harness

LLM-as-Judge eval gates blocking production when precision drops below 0.85. 1,690 versioned ground-truth samples.

0
years shipping production ML & GenAI systems
0
foundation models orchestrated under FacadeDriver
0
versioned ground-truth samples in the LLM evaluation framework
0
concurrent internal devs on the multi-tenant ML platform at Airbnb
0
data consistency across multi-tenant workloads

System Design Architecture

Production architectures I built at Airbnb, mapped to GC AI's product surface

FacadeDriver: Multi-Model LLM Orchestration

30+ foundation models with routing, retry, fallback, and eval-gated accuracy

FacadeDriver: Multi-Model LLM Orchestration
GC AI MappingGC AI Exact Quote System: Same pattern, 5 models tuned per task, verbatim citations on every output
// 03 - Experience

A tour through five production stacks.

Five companies, four industries, one common thread: production-grade systems that data scientists trust.

  1. 2024 - Now
    Airbnb - San Francisco

    Senior Software Engineer - Gen AI Development Experience

    GenAI tooling, Labelbox workflows, PII pipelines, Airflow. 30+ LLMs, 128+ concurrent internal devs, 10K rows/run.

    Contract
  2. 2024
    Eli Lilly - Philadelphia

    Senior Software Engineer - Dose Management System

    Dose order management for radiopharmaceuticals on KUBED. Spring Boot, React, Temporal, Istio. 99.9% uptime.

    7 mos
  3. 2023 - 24
    Southwest Airlines - Dallas

    Senior Software Engineer / Data Scientist

    Kafka streams for real-time flight tracking. RabbitMQ microservices. 95% test coverage via Gradle / Java port.

    1 yr 1 mo
  4. 2021 - 22
    Shell PLC - Houston

    Senior Software Engineer / Data Scientist

    NLP text-classification for drilling-loss events. CNN autoencoders for acoustic impedance. Responsible-AI POC with LIME / SHAP.

    1 yr 11 mos
  5. 2017 - 19
    Oracle India - Bangalore

    Software Engineer

    Cloud HCM integrations, 50K+ LOC. ERP analytics for 13 business units. ELK + Kinesis fraud-detection POCs.

    2 yrs

The stackI've shipped with.

Tools I'd reach for tomorrow, ordered by how often they ship in production at my last 5 jobs.

PythonLLMs / GenAIPyTorchAirflowFlaskFastAPIStreamlitSQLAlchemyCeleryRedisPresidio (PII)LabelboxHive · TrinoSpark · PySparkKafkaDockerKubernetesAWSAzureOpenTelemetryDatadogPythonLLMs / GenAIPyTorchAirflowFlaskFastAPIStreamlitSQLAlchemy
Spring BootJavaReactTypeScriptRabbitMQArgoCDTemporalIstioCrossplaneEntra IDGrafana · LokiSHAP · LIMECNN / AutoencodersNLPDatabricksVector DBsSpring BootJavaReactTypeScriptRabbitMQArgoCDTemporalIstio

What colleagues say

He doesn't just deliver - he continuously looks for ways to make things better.

Sai has been an outstanding partner in the deployment of the BPIVA tool, and I want to take a moment to recognize his incredible contributions. Thanks to Sai's efforts, the BPIVA tool has had a significant impact on reducing non-value-added work, enabling the BPI team to shift their focus to high-impact, actionable tasks exactly where their energy should be. What truly sets Sai apart is his deep understanding of technology combined with his ability to quickly grasp tool requirements and translate them into real solutions. He doesn't just deliver, he continuously looks for ways to enhance and upgrade the tool's capabilities, ensuring it evolves alongside our team's needs. None of this would have been possible without Sai's dedication and expertise. He is a truly great partner who consistently goes above and beyond to deliver excellence. Thank you, Sai, for everything you do! 🙌

AS

Ameet Shinde

Senior Manager, BPI · Airbnb

Letter to GC AI hiring team

GC AI is building
what I've already built.

Your Applied AI team is building multi-model LLM orchestration (5 models tuned per task), legal-grade evals that harden prompts and retrieval, RAG pipelines for complex document analysis, and rapid prototyping frameworks for weekly model releases. That is exactly what I do at Airbnb today. I own the BPI Virtual Analyst - a GenAI platform abstracting 30+ foundation models behind FacadeDriver with routing, retry, fallback, and graceful degradation, serving 128+ users across 4 partner teams. I built the 23-version LLM evaluation harness with 1,690 ground-truth samples and LLM-as-Judge gates that block production when precision drops below 0.85. I built end-to-end RAG pipelines with reranking and contextual compression, multi-agent systems with tool calling and agentic memory, and rapid prototyping frameworks consumed by other teams. Your Exact Quote system demands verbatim accuracy on every citation. My eval gates demand the same. I want to bring this to GC AI.

// 04 - Schooling

Where I learned to do it.

Graduate
M.S. Computer Science - Machine Learning and AI focus
NYU Tandon - Top Coursework: ML Systems, NLP, Distributed Systems
Certifications & Writing
Open source & Medium
Contributes to LiteLLM, LangChain, LiveKit - Writes on Medium about LLM eval tooling

Let's talk
GC AI.

Open to Applied AI roles at GC AI

SE 4/5 - Remote (Houston, TX) - H-1B transfer (I-140 approved, AC21 portable). Available immediately.

Start a conversation
Multi-tenant ML platformsFoundation modelsLLM evaluationAgentic AI frameworks