- Built a conversational AI agent for Airbnb's internal prompt-engineering platform, using retrieval-augmented generation and a confirm-before-acting human-in-the-loop workflow, designed with the platform team.
- Instrumented distributed tracing with OpenTelemetry across the agent's inference pipeline, adding mesh-level proxy authorization and span propagation to restore observability for production debugging and monitoring.
- Built a structured intake pipeline and a mode-aware generation system with real-time diff rendering, using few-shot grounding over a curated production dataset to draft prompts for input types it hadn't seen.
- Built an evaluation harness on the platform's LLM execution runtime, with persistent result caching and a curated fallback dataset, to validate agent outputs across seven task categories.
Timeline
Drawn to scale · Sep 2023 to May 2027
Experience
3 internships
- Shipped an MVP iOS app from scratch in Swift with an MVVM architecture, REST APIs and Firebase real-time sync, covering 15+ core features.
- Reached an 85% task completion rate and grew daily active users 25% through A/B-tested UX iterations.
- Connected the app to backend infrastructure through a Chrome extension, holding 99.5% uptime.
- Designed a PostgreSQL schema and PgBouncer connection-pooling layer for concurrent trading-account requests at peak market hours, cutting query latency 40% and raising throughput 25%.
- Built a REST integration with the Interactive Brokers and Twilio APIs for automated account provisioning.
- Added automated tests and CI/CD pipelines on GitHub Actions, cutting deployment errors 35%.
Projects
Research, a hackathon, and a long build
Parameter Golf
OpenAI Model Craft Challenge · independent research
Apr–May 2026 · PyTorch, CUDA, C
Train the best language model that fits in 16 MB, then evaluate it on FineWeb inside a 600-second budget. The score is bits per byte (BPB), and lower is better.
- Best score
- 1.0567 BPB
- Size cap
- 16 MB
- Eval budget
- 600 s
- Hardware
- 8×H100
- Ablations
- 60+
- Stacked a token-only n-gram tilt, asymmetric logit rescaling and phased test-time training on a 16 MB model. The submission briefly held the top spot on the open leaderboard at 1.0567 BPB across three seeds.
- Ported a C n-gram state machine and its Python compute path from an earlier merged PR, then moved precompute inside the eval timer so a 167 s hint table plus 355 s of phased TTT fit the 600 s budget on all three seeds.
- Ran 60+ tracked ablations across architecture, quantization (GPTQ with LQER asymmetric rank-4) and TTT settings. One setting,
PHASED_TTT_NUM_PHASES=1, cut eval time by 49 s and improved BPB on the same hardware. - Read through 20+ merged competition PRs to map the architecture lineage, then built on the latest baseline while keeping the standard CaseOps tokenizer and BPB method.
Relay
MongoDB NYC hackathon · team of 3
Sep 26, 2026 · Python, MongoDB Atlas, ElevenLabs
An agent that turns disaster-recovery conversations into tracked follow-ups over six weeks. MongoDB is the source of truth, so the agent can crash and resume without losing anything. I owned the voice calls, the safety layer and the evaluation.
- Built in
- 1 day
- Lines added
- ~8,600
- Tests
- 105
- Emergency phrases
- 60+
- Languages
- 3
- Connected ElevenLabs voice calls to the agent through a custom LLM endpoint that streams a short filler line first, then answers with a fast model, so replies stay quick on a live call.
- Wrote a safety layer that checks every utterance against 60+ emergency phrases in English, Spanish and Mandarin before any model sees it. A match plays a fixed 911 message and alerts a person on the team.
- Built the evaluation harness that compares Relay with a plain long-context LLM under identical conditions, with deterministic scoring that gives credit only when a household is flagged before its deadline.
StoryForge
AI interactive fiction engine · personal project
Nov 2025–Apr 2026 · Python, FastAPI, React, TypeScript
Narrator, Character and Director agents build branching stories that remember what happened, backed by vector, graph and relational stores.
- Agents
- 3
- Story paths
- 50+
- Datastores
- 3
- Interactions
- 100+ / session
- Orchestrated Narrator, Character and Director agents to generate branching narratives across 50+ dynamically generated story paths.
- Built layered memory: ChromaDB for vector recall and Neo4j for a causal graph of story events, so characters remember and consequences carry forward.
- Built the full stack: FastAPI with JWT auth, PostgreSQL for story state, WebSockets for multi-user sessions, and a React/TypeScript client with real-time scene generation.
Toolbox
What I build with
Languages
Go, Python, C/C++, Java, TypeScript, JavaScript, Lua, Bash, SQL
Systems
Distributed systems, concurrency, lock-free data structures, gRPC, real-time messaging, WebSockets, Redis
Data and infrastructure
PostgreSQL, PgBouncer, MongoDB, Airflow, Docker, Kubernetes, AWS, GCP, Jenkins, OpenTelemetry, Linux, Git
ML
PyTorch, CUDA, distributed training, vLLM, LangGraph, retrieval-augmented generation
- B.S. in Computer Science, with an additional major in Business Administration (AI concentration).
- Coursework: Parallel and Sequential Data Structures and Algorithms, Probability and Computing, Introduction to Machine Learning, Functional Programming, Principles of Imperative Computation, Matrices and Linear Transformations.
Contact
Fastest reply by email