Data Engineer II at Thinking Machines Data Science
Builds AI agents, data pipelines, and cloud infrastructure for enterprise clients across financial services, investment management, education, and compliance.
- Led the data products workstream in a year-long engagement, partnering with C-suite leaders, directors, and business units to define and ship priority data products on Azure Databricks.
- Designed and productionized a daily Single Customer View pipeline consolidating ~15 million records from four enterprise systems into ~7 million unique customer keys used across 10+ business units.
- Built a probabilistic record-linkage engine with Splink that surfaced 748,000 candidate duplicate pairs missed by exact matching, with a five-tier confidence framework validated at >99.9% accuracy.
- Cut Single Customer View runtime from over 4 hours to under 1 hour and co-designed a Next Best Product pipeline covering 12 product types and four customer segments.
- 15M → 7M
- records consolidated to unique customer keys
- 748,000
- duplicate pairs surfaced by record linkage
- >99.9%
- five-tier confidence accuracy
- 4h → <1h
- pipeline runtime
Investment Data Platform
- Owned backend feature development across 10+ interconnected microservices powering deal discovery, due diligence, and portfolio monitoring workflows.
- Built backend services and data pipelines with FastAPI, Dagster, Kubernetes, and Snowflake, processing terabytes of vendor data from Sustainalytics, PitchBook, Bloomberg, and MSCI.
- Improved platform reliability and observability through production monitoring on Kibana and Grafana.
Enterprise Document Intelligence Platform
- Iterated on Snowflake Cortex AI agent workflows with the Knowledge Graph team so investment officers could extract data from long-form documents, run contextual queries, and analyze portfolios at scale.
- Owned agent tooling, prompt engineering, evaluation, workflow tuning, and production incident triage across the document-intelligence stack.
- Partnered with a ~15-person cross-functional team spanning frontend, backend, LLM workflows, MLOps, DevOps, and QA.
- Spearheaded the infrastructure and ingestion workstream for a student-at-risk prediction platform on Azure Databricks.
- Provisioned platform infrastructure with Terraform, built API ingestion across three source systems, and migrated Excel-based linear-regression models into production Databricks jobs.
- Accelerated delivery by integrating AI coding agents into the workflow with custom agent skills, hooks, and command layers.
- Led the ingestion and transformation workstream, building end-to-end pipelines from four source systems in SQL Server and MongoDB into BigQuery.
- Designed surrogate-key strategies to resolve identifier collisions between member and visitor systems and modeled production reporting tables for attendance analytics.
- Resolved memory failures in legacy Airflow DAGs with chunked CSV processing before production rollout.
- Designed and delivered a six-course enablement curriculum covering Python and Power BI from beginner to advanced in four days.
- Authored a standardized data-ingestion scoping framework RFC for the engineering consulting team.
- Created post-exam study guides that supported certification prep across the company.
- Built a proof-of-concept integration connecting ChatGPT to Databricks through an MCP server.


