Data Science @ UC San Diego · Class of 2028

TijilChhabra

I build ML systems, agent pipelines, and data products. Three internships in, the ones I'm proudest of are the ones that refuse to make things up.

01 / shipped output

Receipts, not vibes.

Lawgical · Document AI + Claude
+6–10pts

OCR accuracy from an LLM post-processing stage, on top of a 90%+ baseline across 6 document types

Lawgical · Next.js + Express
40–50%

less manual data entry after shipping a full-stack OCR service to production

Lawgical · Claude Opus
0+

US visa document types auto-labeled by a classification service, replacing manual sorting

PwC · LangGraph + Postgres
0

nudges on a multi-agent BFSI platform, scaled from 10 through 7 reusable capability packs

PwC · Python + RAG
0

hallucinated products from a grounded recommender with fail-closed eligibility and 35/35 tests passing

offline eval · demo customer dataset
DGLiger · Databricks, FAISS, Prophet
0

data services in one internship: bank ETL pipelines, vector-similarity targeting, and cashflow forecasting

Portfolio Manager · FastAPI + Postgres
0

passing tests across a FIFO ledger, live pricing, and LangGraph agents

Recipe Ratings · scikit-learn
0→12%

minority-class recall from a class-balanced Random Forest tuned with 5-fold GridSearchCV

02 / experience

Where I've shipped.

PwC

SWE Intern (AI/ML)
Jun 2026 – Aug 2026 · New Delhi
  • Built a multi-agent nudge platform for a top-3 Indian private bank, with a BFSI admin surface and a customer surface, where each agent owns a distinct task.
  • Orchestrator / worker / validator pattern on a LangGraph state machine and Postgres schema, with rules-based arbitration, consent and fatigue gating, human escalation, and a reason code on every outcome.
  • Scaled the nudge catalogue from 10 to 33 through 7 reusable capability packs, across banking, lending, insurance, and wealth.
  • Replaced LLM-generated product suggestions, which invented offers that didn't exist, with a grounded Python recommendation engine: fail-closed eligibility, RAG scoped to explanation only, and product ids validated in code with a deterministic fallback.
  • Added 20 unit tests (suite green at 35/35) with zero new runtime dependencies.
LangGraphPythonPostgreSQLClaude APIRAG

Lawgical

Software Engineer Intern (AI/ML)
Jul 2025 – Sep 2025 · Remote
  • Shipped a full-stack document OCR service (Next.js + Express) into the production website, cutting manual data entry 40–50%.
  • Built OCR pipelines on GCP Document AI for 6 document types at 90%+ accuracy, then diagnosed recurring extraction errors and added an LLM post-processing stage for a further 6–10%.
  • Built a document classification service on Claude Opus that auto-labels 10+ US visa document types, replacing manual sorting.
Next.jsExpressGCP Document AIClaude Opus

DGLiger

Data Science Intern
Jul 2024 – Sep 2024 · New Delhi
  • Three services in one internship, spanning ingestion, modeling, and KPI definition, for a partner bank's Data-as-a-Service initiative.
  • Designed schemas and built distributed ETL pipelines on Databricks, turning raw transactional data into analytics-ready tables for dashboards and downstream models.
  • Built a FAISS vector-similarity targeting service to surface high-value prospects for merchant partners.
  • Built an SME cashflow forecasting module with Prophet for short-term liquidity prediction that informed merchant credit decisions.
DatabricksPySparkFAISSProphet

DataHacks

Director · DS3 @ UCSD
Oct 2024 – Present · La Jolla
  • Directed San Diego's largest hackathon, a 36-hour, MLH-certified event: 1000+ registrations and 450 hackers from 70+ universities.
  • Secured $90K in sponsorship from companies including Nvidia, Amazon, Google, and Qualcomm through data-backed pitches and ROI reporting.
  • Owned end-to-end operations, budgeting, and stakeholder communication across logistics, sponsorship, and judging.
03 / projects

Things I built for fun
(and a grade).

★ featured · live

Portfolio Manager

One dashboard for your whole portfolio across Indian NSE and US markets. Drop in an Excel or broker CSV and get live prices, true cost basis, risk alerts, and an AI analyst that reads the market news for you.

  • Live prices from Yahoo Finance (yfinance), batched across NSE and US tickers, cached for 15 minutes and refreshable on demand.
  • News merged from Yahoo Finance and Google News RSS for every holding.
  • Fundamentals for any ticker from Yahoo Finance, with Finnhub as a fallback for US ratios.
  • FIFO cost-basis ledger with realised and unrealised P&L, plus alerts for concentration, small-cap exposure above 5%, and allocation drift.
  • Weekly AI briefing: LangGraph agents (orchestrator → market → sector specialists) mark each development as positive or negative for your holdings.
  • Model tiering: Claude on premium, Gemini or Groq on the free plan, all behind one analyst interface.
  • Supabase Auth for sign-up and login, and Supabase Postgres where every row is scoped to your account.
  • Deployed on Render, with a daily GitHub Actions job that refreshes prices and snapshots.
FastAPISupabasePostgreSQLLangGraphClaudeGeminiGroqyfinanceFinnhubRenderGitHub Actions
hackathon · biophysics

T Cell Swarming Simulation

Why do T cell swarms stop themselves? I wrote the simulation engine for our 4-person team at UCSD's Active Matter hackathon, in under 24 hours.

  • ABP self-propulsion, Vicsek alignment, chemotaxis, receptor desensitization
  • Limit tests switch off each mechanism and assert collapse to the known simpler regime
PythonNumPyMatplotlib
nlp · deep learning

TweetGuard

Hate speech detection on 31,962 labeled tweets where only 7% are positive, so a 0.5 threshold is the wrong default.

  • NLTK preprocessing → GloVe 50d embeddings → 3 stacked BiLSTMs
  • Decision threshold tuned off the precision-recall curve for the minority class
TensorFlowKerasNLTKGloVe
stats · DSC 80

Do High-Calorie Recipes Taste Better?

83K recipes and 732K Food.com reviews. Short answer: no (p = 0.072).

  • Permutation tests and MNAR missingness analysis
  • Class-balanced Random Forest, 5-fold GridSearchCV, lifting minority-class recall from 0% to 12%
pandasscikit-learn
full-stack · scraping

SoleSavings

A sneaker price aggregator that pulls live listings across marketplaces into one SQLite store for cross-site comparison.

  • Plugin-style scrapers, so new marketplaces don't touch core code
  • Automated ingestion into a schema built for fast historical price queries
FlaskSQLiteBeautifulSoupSelenium
04 / the graph

Everything, connected.

I built this with graphify from my career database, using only verified facts: . Hover to trace, click a node to jump to it, drag to play.Tap a node to trace it, tap again to jump there, pinch to zoom. Dashed lines are inferred links.

05 / stack

The toolbox.

// languages

Python · SQL · JavaScript · Java · R · HTML/CSS

// ai & ml

LangGraph · Claude API · RAG · FAISS · TensorFlow / Keras · scikit-learn · Prophet · NLTK · GloVe

// data

pandas · NumPy · PySpark · Databricks · PostgreSQL · SQLite · SQLAlchemy · D3.js

// backend & infra

FastAPI · Flask · Next.js · Express · React · Docker · GCP · Supabase · Render · Git

06 / education

Still learning.

Sep 2024 – Jun 2028

UC San Diego

Bachelor's, Data Science

  • DSC 30 · Data Structures & Algorithms
  • DSC 40A/B · Theoretical Foundations of Data Science
  • DSC 80 · Practice & Application of Data Science
  • DSC 100 · Database Systems · DSC 106 · Data Visualization
  • MATH 183 · Statistical Methods
certification

AWS Certified Cloud Practitioner

Amazon Web Services Training and Certification

07 / contact

Let's build something.

tchhabra@ucsd.edu
github ↗ linkedin ↗
↑↓ navigate↵ selectesc close