Kyle Naranjo06 projects · 2022 – Present

Projects

The systems behind the résumé lines — enterprise data platforms, agentic AI, and the research that started it all. Clients are anonymized; the numbers are exact.

Single Customer View

A year-long engagement with a major Philippine bank, partnering with C-suite leaders and business units to ship priority data products on Azure Databricks. The centerpiece is a daily Single Customer View pipeline consolidating ~15 million records from four enterprise systems into ~7 million unique customer keys used across 10+ business units.

Exact matching misses people. A probabilistic record-linkage engine built with Splink surfaced 748,000 candidate duplicate pairs the exact-match logic could not see, scored through a five-tier confidence framework validated at >99.9% accuracy — while pipeline runtime dropped from over 4 hours to under 1 hour.

15M → 7M
records consolidated to unique customer keys
748,000
duplicate pairs surfaced by record linkage
>99.9%
five-tier confidence accuracy
4h → <1h
pipeline runtime

Investment Data Platform

Owned backend feature development across 10+ interconnected microservices powering deal discovery, due diligence, and portfolio monitoring workflows for a large Singaporean investment holding company.

The platform processes terabytes of vendor data from Sustainalytics, PitchBook, Bloomberg, and MSCI through FastAPI services and Dagster pipelines on Kubernetes, landing in Snowflake — with reliability and observability handled through production monitoring on Kibana and Grafana.

10+
interconnected microservices
4
vendor data feeds processed

Enterprise Document Intelligence Platform

Iterated on Snowflake Cortex AI agent workflows with the Knowledge Graph team so investment officers could extract data from long-form documents, run contextual queries, and analyze portfolios at scale.

Owned agent tooling, prompt engineering, evaluation, workflow tuning, and production incident triage across the document-intelligence stack, partnering with a ~15-person cross-functional team spanning frontend, backend, LLM workflows, MLOps, DevOps, and QA.

~15
person cross-functional team

Student-at-Risk Prediction Platform

Spearheaded the infrastructure and ingestion workstream for a student-at-risk prediction platform on Azure Databricks: platform infrastructure provisioned with Terraform, API ingestion built across three source systems, and Excel-based linear-regression models migrated into production Databricks jobs.

Delivery was accelerated by integrating AI coding agents into the workflow with custom agent skills, hooks, and command layers.

3
source systems ingested via API

ChatGPT-to-Databricks MCP Server

A proof-of-concept integration connecting ChatGPT to Databricks through a Model Context Protocol server — the experiment behind the MCP spec's place in the resources list.

The PoC demonstrated a governed path from a conversational assistant to warehouse queries without another one-off API wrapper, and now feeds Kyle's talks on agentic workflows.

AIComprehend

Co-authored and published the IEEE AIComprehend paper, then built and deployed the full-stack Django application used in the study itself.

A four-week controlled study with 58 high school students showed a 13.9% improvement in test scores — research that started the thread running through the rest of this list.

13.9%
test-score improvement
58
students in the controlled study
4 wk
study duration

Have a data or AI problem worth building for?

Get in Touch