Jay Shah
Senior Data Scientist building production AI applications and foundational model systems with robust evaluation at scale.
About
Senior Data Scientist at 6sense building production AI applications, foundational model systems, and agent evaluation frameworks. Expert in model explainability, Python, AWS, and GCP with cross-team leadership from roadmap to production.
Work Experience
6senseSan Francisco, CA
Senior Data Scientist
AvathonPleasanton, California
Data Scientist III
AvathonSunnyvale, California
Data Scientist II
Avathon (Acquired Ensemble Energy)Palo Alto, California
Data Scientist
Avathon (Acquired Ensemble Energy)Palo Alto, California
Data Science Intern
Texas A&M UniversityCollege Station, Texas
Graduate Research Assistant
Utilities and Energy ServicesCollege Station, Texas
Student Analyst
DataKindSan Francisco, California
Data Ambassador
Education
Texas A&M University
Gujarat Technological University
Skills
AI & Foundational Systems
Libraries & ML Tools
MLOps & Data
Software Engineering
Projects
Text Watermarking Lab
Experiments with keyed text watermarking, statistical detection, model behavior, calibration, and the effect of editing on the signal.
AgentEval Suite — Specialized Evals for Production Agents
Scenario-driven evaluation harness for domain agents with synthetic task generation, tool-usage tracing, success and latency metrics, and judge models for regression testing.
Pi Agent Extensions
A collection of extensions and themes for the Pi coding agent, including sessions, structured questions, handoffs, multi-agent workflows, and review tools.
Session Aggregator
A local tool for syncing, searching, and exporting AI coding sessions across development tools, with a terminal UI and semantic search.
Medha IDE
A local-first SQL IDE for flat files that combines DuckDB, FastAPI, Vite, and LangGraph for semantic query generation.
Arka — Config-Driven Synthetic Data Generation
A config-driven pipeline for generating and filtering fine-tuning data with multi-source ingestion, Evol-Instruct, MinHash and LSH deduplication, and SQLite checkpoints.
Humanizer-RL — AI Text Humanness Scorer
A text humanness scorer and reinforcement-learning pipeline that turns model evaluations into a reward function for Gemma fine-tuning.
Entity Resolution POC
An entity-resolution evaluation comparing dense embeddings with BM25 and Matryoshka Representation Learning for efficient retrieval at scale.
Pravāha — AI Search Engine
A local search assistant that combines web search, document retrieval, agent tools, and language models with specialized ranking and chunking.
Gujarati Llama
A Llama 2 7B model fine-tuned on 60,000 bilingual English-Gujarati pairs for low-resource language use cases.
StreamLens
A multi-model retrieval-augmented system for interacting with autonomous-vehicle video streams, built for the LlamaIndex RAG-a-thon.
NeuroBuddy
A personalized chatbot for mental-health support and resources, built with Mistral AI and Whisper models during a hackathon.
Power Curve Estimation for Wind Energy Farms
Statistical and machine-learning models for estimating wind-farm power curves and supporting energy-system optimization.
Customer Relationship Prediction
Classification models for predicting churn, appetency, and up-selling behavior for a mobile network operator.
Phase 1 Analysis
Multivariate quality-control analysis for an industrial forging process using principal components and control charts.