Text-to-SQL Benchmark & Retrieval Tuning
Benchmarked a text-to-SQL pilot on GPT-4o + pgvector across 150 gold-SQL pairs, then rewrote schema docs into the retrieval context to lift execution accuracy.
Data Analyst & AI Engineer · Toronto
I build analytics and the AI systems behind it — SQL & BI, ETL, RAG, text-to-SQL, and LLM evaluation. Fewer claims, more receipts.
UofT — Math (Stats & English minor), Dec 2026 · open to new-grad roles
Benchmarked a text-to-SQL pilot on GPT-4o + pgvector across 150 gold-SQL pairs, then rewrote schema docs into the retrieval context to lift execution accuracy.
Replaced a 4-hour weekly Excel report with 6 validated Power BI views over a Snowflake warehouse — one source of truth for the account team.
Designed the retrieval layer over ~400 articles with LlamaIndex + Chroma; selected the chunking strategy that scored highest on a 120-question benchmark.
Built the SQL feature pipeline and time-based train/test split behind a client’s first predictive model over a 10M-row warehouse, held on the out-of-time set.
I’m a University of Toronto math student (statistics and English minors) who ended up equally at home in a SQL console and an eval harness. I like the part of the job where a fuzzy business question becomes a number, and the number becomes a decision.
Most recently at Avocette I worked across both sides of that line — owning weekly BI reporting and dashboards, and building applied-AI systems: a text-to-SQL benchmark, a RAG retrieval layer, an LLM evaluation rubric, and a client’s first predictive model. Before that I ran a profitable Shopify brand end to end, which is where I learned to respect a P&L.