# Jordan Kail > Staff Software Engineer at Together AI. With over 13 years of experience in AI, analytics and machine learning, I now lead the agents platform and data engineering at Together AI, building the agent harness, internal agents factory and proprietary IP behind them. Before that I built AI classifiers at Meta and model-governance frameworks at Prove, and turned data into impact for startups and Fortune 50 companies alike. Based in Denver, CO, United States. This file is generated from the same data as the website; the site is a single-page app, so use these plain-text and JSON renditions instead of scraping it. ## About I'm a software engineer with a deep curiosity for technology — always exploring AI, machine learning, and data systems. I've helped build efficient, scalable platforms, complex pipelines, and production transformer-model workflows. I love collaborating with others to create technology that makes a real impact. When I'm not working, you'll likely find me building computers, hiking with my dog in the mountains, or experimenting with new ML tools. ## Experience ### Staff Software Engineer, Together AI 02/2025 - Present | San Francisco, CA | https://www.together.ai/ Together AI is an AI acceleration cloud offering fast inference, fine-tuning, and dedicated GPU clusters for open-source and custom models, backed by a research team that builds and open-sources frontier models and systems optimizations. - Tech lead for data engineering since joining: grew the team from just myself to 15+ engineers while setting technical direction for the data platform behind Together's AI acceleration cloud. - Tech lead for the agents platform: created the agent harness, built an internal agents factory and used it to deliver production agents, and contributed to multiple patent-pending innovations in agent systems. - Built and operate agents for infrastructure automation, finance automation and other internal workflows that augment the teams they serve, extending team capacity while improving on-call response and infrastructure maintenance. - Build the data platform behind Together's AI acceleration cloud: the pipelines, storage, and telemetry that turn inference and training traffic into product, reliability, and capacity signals. - Build tracing, replay and evaluation harnesses so internal teams can measure and regression-test LLM agents before they ship. - Build tracing, replay, and regression-testing harnesses for LLM agents, giving internal teams a repeatable way to evaluate agentic features before they ship. - Own core data platform services for inference, fine-tuning, and dedicated GPU cluster workloads, making platform usage measurable end to end. - Develop streaming and batch pipelines that make model-serving and GPU-fleet telemetry queryable in near real time for capacity planning and reliability engineering. - Design data infrastructure for large-scale training and inference workloads, covering dataset curation, lineage, and quality controls for open-model work. - Partner with research, infrastructure, and product teams to standardize data contracts and lineage for training-data and evaluation workflows. - Mentor engineers and set technical direction for the data platform and agents platform architecture. Technologies: Python, SQL, Go, TypeScript, Kubernetes, PyTorch, agent-platforms, agent-harnesses, llm-evaluation ### Staff Software Engineer - Data, Prove Identity 06/2023 - 01/2025 | Denver, CO | https://www.prove.com/ Prove is the modern platform for consumer identity verification. We enable businesses to securely verify customer identities in real-time, with the highest level of compliance and user experience. - Built agents and harnesses that govern the statistical models used for identity and fraud resolution; these became core frameworks. - Built AI-driven Retrieval-Augmented Generation (RAG) chatbots with Airflow, LangChain, and OpenAI, automating 150+ human-hours weekly. - Engineered AI-powered RAG chatbots leveraging Airflow, LangChain, and OpenAI APIs to automate customer service workflows, reducing manual efforts by 150 hours per week. Addressed NLP challenges in ambiguous query handling, driving a 20% improvement in first-contact resolution. - Spearheaded a company-wide migration from legacy Java/Oracle infrastructure to a cloud-native Go/PostgreSQL stack, cutting API response times from 30s to 12ms and slashing operational costs by 95%. Overcame challenges with service downtime, seamlessly integrating AWS services (Athena, S3, EC2) for enhanced scalability. - Managed and mentored 9 data engineers and 6 data scientists, driving a 20% improvement in project delivery timelines. Collaborated with the VP of Platform Engineering on critical initiatives, serving as the primary data liaison to ensure seamless cross-departmental coordination. - Deployed an event-driven data streaming platform with Go, Kafka, and Flink, transitioning 1,200+ batch jobs to real-time processing. Reduced data availability lag from days to seconds, significantly boosting dashboard accuracy and enabling data-driven decision-making for executive teams. - Implemented real-time telemetry services with Prometheus, Grafana, Splunk, and AWS CloudWatch, addressing complex monitoring needs. Reduced incident response times by 40% while maintaining 99.9% system uptime for mission-critical services. - Designed and developed data pipelines using Spark, Airflow, and DBT to streamline ETL processes, improving performance by 35%. Reduced report generation timelines from days to hours, enhancing reporting capabilities for product managers and business teams. - Established robust data governance frameworks with Apache Atlas, ensuring full GDPR and SOC2 compliance. Implemented metadata standards, lineage tracking, and access policies, reducing audit preparation times by 30%. - Optimized CI/CD pipelines for Apache Beam and Kubernetes deployments, resolving bottlenecks and reducing deployment times by 30%. Improved release frequency to support faster feature rollouts and iterative development cycles. - Migrated computationally intensive SQL queries from RDS to Spark DataFrames, improving query performance by 70% and reducing compute costs by 30%. Enabled seamless processing of large datasets for business-critical applications. Technologies: Python, SQL, Go, JavaScript, TypeScript, Java, AWS, model-governance, RAG ### Senior Data Engineer, Meta | Facebook 01/2021 - 09/2022 | Seattle, WA + Menlo Park, CA + Remote, USA | https://about.fb.com/ Meta Platforms, Inc. is an American multinational technology conglomerate based in Menlo Park, California. It was founded by Mark Zuckerberg, along with his college roommates and fellow Harvard University students Eduardo Saverin, Andrew McCollum, Dustin Moskovitz and Chris Hughes, originally as TheFacebook.com—today's Facebook, a popular global social networking service. - Built AI/ML classifiers and the data, feature and evaluation pipelines around them. - Designed and deployed 100+ ML pipelines using Airflow, Spark, and PyTorch, integrating NLP and computer vision models. Improved ad targeting precision and personalized notifications, driving higher engagement across billions of users. - Lead data engineer for Facebook Public Groups, Community Chats, and cross-platform initiatives, leading 10+ engineers. Collaborated with data science, ML, and hardware teams to align product goals, resulting in a 20% faster delivery of cross-platform features. - Designed exabyte scale data models for Community Messenger to handle multi-platform data streams, increasing user engagement by 35% across Instagram, WhatsApp, Facebook, and Quest despite complex cross-platform dependencies. - Developed modular frameworks for data pipelines, streamlining data flows for thousands of engineers. Reduced integration issues by 40% and accelerated feature deployments by 25% through automation and standardization. - Built telemetry systems capable of processing 10M+ events/second, improving signal quality by 20%. Developed Jinja-based monitoring tools, reducing downtime by 15% and ensuring reliable system performance. - Implemented graph and entity models to support 10 billion monthly interactions, ensuring seamless experiences across Instagram, WhatsApp, Facebook, and Quest. Addressed latency and consistency challenges to maintain real-time performance. - Served as the liaison between Facebook Messenger, Groups, and the early precursor to the LLaMA project, fine-tuning models for automated group management. Increased group engagement by 15% in pilot testing, paving the way for future LLM-powered features. - Created from scratch QR code group invites, increasing join rates by 50%. Enabled offline engagement in multilingual regions, facilitating family reconnections and shelter logistics coordination during humanitarian efforts. Featured on Tech Crunch. - Designed KPI dashboards to monitor DAUs, MAUs, and engagement trends using tools such as Tableau and internal data visualization frameworks. Enabled real-time insights, increasing engagement by 10% and retention by 12%. - Engineered an automated framework for generating thousands of asynchronous Spark data pipelines, increasing compute efficiency by 66%. Overcame orchestration challenges, improving resource utilization and processing times. Technologies: Python, SQL, JavaScript, TypeScript, PHP, Scala, ai-classifiers, PyTorch ### Consultant - AI & Advanced Analytics, Deloitte 12/2018 - 12/2020 | Atlanta, GA + Menlo Park, CA | https://www2.deloitte.com/ch/en/pages/strategy-operations/solutions/analytics-and-cognitive.html Deloitte Touche Tohmatsu Limited, commonly referred to as Deloitte, is a multinational professional services network. Deloitte is one of the Big Four accounting organizations and the largest professional services network in the world by revenue and number of professionals, with headquarters in London, United Kingdom - Automated approximately 31% of human processed healthcare claims using transformer machine learning models, saving 250,000+ hours annually. - Optimized exabyte-scale video reliability metrics, reducing daily processing time by 90% while expanding metric coverage. - Automated 31% of healthcare claims processing with transformer-based ML models, addressing data inconsistencies and regulatory constraints. Saved 250,000+ human-hours annually and improved claim accuracy by 20% - Optimized exabyte-scale video reliability metrics using distributed data processing frameworks. Cut daily processing time by 90% and expanded metric coverage across multiple product lines, improving monitoring precision. - Led teams of 20+ consultants on high-profile Fortune 50 engagements, delivering advanced technical solutions aligned with client needs. Achieved 100% on-time project delivery across multiple engagements. - Developed a risk detection service using machine learning algorithms to flag high-value insurance accounts. Reduced billing errors by 30% and improved revenue recovery through early detection of anomalies. - Engineered NLP models inspired by Google's Transformer architecture within six months of its release. Reduced billing errors by 20% by implementing cutting-edge language models for document processing. - Implemented massively parallel data pipelines using Spark and asynchronous frameworks, reducing healthcare claim turnaround times by 40%. Improved processing efficiency for high-volume workloads. - Re-architected live video infrastructure to align with emerging short-form content trends, saving $100 million and 14 months of development time. Ensured seamless adoption of new formats across the platform. - Optimized caching strategies, leveraging Redis and CDN-layer optimizations to cut costs by $10 million annually. Enhanced content delivery efficiency and reduced latency for high-traffic web services. - Built predictive analytics models for server uptime using time-series forecasting techniques, improving video delivery reliability by 30% and enhancing user experience through proactive maintenance. - Revamped A/B testing frameworks for billions of daily users, resolving data inconsistencies and enabling more precise feature rollouts. Accelerated data-driven decision-making with improved statistical significance tracking. - Leveraged Python, SQL, Java, TensorFlow, PyTorch, Spark, Airflow, Docker, and Kubernetes to develop and deploy scalable technical solutions. Delivered projects across machine learning, real-time analytics, and data pipeline automation for Fortune 50 clients. Technologies: Python, SQL, JavaScript, TypeScript, PHP, Rust, AWS, Google Cloud ### Senior Data Engineer, Wide Open West 11/2017 - 12/2018 | Denver, CO | https://www.wowway.com/ Wide Open West is the sixth largest cable operator in the United States. The company offers landline telephone, Cable Television, and broadband Internet services - Built Machine Learning applications using custom classification and churn models, driving a 22% YoY increase in customer package upgrades. - Led a team of 5 data practitioners, providing BI and data insights to sales, product, and engineering teams company-wide. - Built machine learning models, including custom classification and churn prediction algorithms (e.g., logistic regression, K-means clustering). Increased customer package upgrades by 22% YoY and reduced churn by 15%. - Managed a team of 5 data practitioners, providing BI and insights to sales, product, and engineering teams. Established KPIs and standardized reporting through governance committees, driving a 10% increase in sales performance. - Developed a Kafka-powered real-time analytics platform, streaming data from field technicians and delivering instant job updates via a custom web portal. Reduced service completion times by 40%, replacing 20-minute phone calls with real-time notifications. - Built scalable, cloud-based data solutions leveraging AWS services (SageMaker, S3, Redshift, and Athena). Improved data accessibility and reduced query times by 30% across sales and operations teams. - Automated marketing campaigns using SendGrid to engage at-risk customers, reducing churn by 15%. Implemented multi-channel unsubscribe mechanisms, ensuring 100% compliance with communication preferences. - Delivered geospatial insights using GIS tools to support sales in Arkansas and Alabama. Optimized resource allocation down to the city block level, driving an 18% increase in sales conversions and empowering door-to-door teams. - Developed a dynamic revenue forecasting tool using Python and SQL to set bonus targets and calculate commissions by territory. Increased sales velocity and retention, driving a 12% increase in quarterly sales. - Created a sales funnel dashboard with Tableau to monitor add-on targets, conversions, and installations. Identified millions in unrealized losses, leading to strategic reallocations and improved market performance. - Applied machine learning models and geospatial analytics to optimize network NUC performance, reducing infrastructure build-out costs by 25%. Implemented continuous deployment with GitHub, ensuring code quality through design principles and best practices. - Technologies Used: Python, Java, JavaScript, SQL, Flask, AWS (SageMaker, S3, EC2, Redshift, Athena), Apache Kafka, SendGrid, GIS tools. Technologies: Python, SQL, JavaScript, AWS ### Data Engineer, Common Spirit Health 09/2016 - 11/2017 | Denver, CO | https://www.commonspirit.org/ CommonSpirit Health is a nonprofit, Catholic health system dedicated to advancing health for all people. It was created in February 2019 through the alignment of Catholic Health Initiatives and Dignity Health. CommonSpirit Health is the largest nonprofit health system in the U.S. with more than 1,000 care sites in 21 states. - Architected and delivered new rest APIs and data lakes, improving data processing time for external partner data products from 7 days to 5 minutes. - Developed machine learning pipelines using Python and SQL to forecast hospital procedures, billing, and staffing. Reduced billing turnaround from 14 to 7 days. - Architected and delivered new REST APIs and data lakes, reducing external partner data processing time from 7 days to 5 minutes. Enhanced data accessibility and scalability through efficient data structures. - Managed tier-one vendor data extracts containing patient records, financial data, and ICD codes. Improved extract performance by 40% and ensured 100% HIPAA compliance through encryption and access control measures. - Developed a data quality dashboard using Tableau and Python, enabling real-time failure detection. Reduced issue resolution time by 50% and improved overall data integrity and operational reliability. - Conducted advanced data analysis using Dimensional Fact Models in SMP and MPP environments, improving query performance by 35%. Delivered actionable insights to executives, enhancing decision-making processes. - Developed flexible big data extracts and real-time CDC lakes using AWS Redshift and Kafka. Enabled faster product delivery, reducing time to market by 20%. - Designed complex data models for highly sensitive UII data, employing encryption and role-based access control. Ensured data security while supporting high-stakes analytics and compliance use cases. - Created analytics dashboards using Qlik Sense, Tableau, and Python to track key metrics. Delivered actionable insights to executives, increasing reporting efficiency by 25% and enhancing operational visibility. - Developed optimized storage and compute solutions, reducing third-party vendor data costs by 45%. Earned recognition from the VP of Business Intelligence for faster issue resolution and improved efficiency. - Migrated data from relational to columnar formats (e.g., Parquet) using Redshift, improving query speed by 40% and enabling large-scale data processing for analytics. - Developed machine learning pipelines using Python and SQL to forecast hospital procedures, billing, and staffing. Reduced billing turnaround from 14 to 5 days, optimized surgical room usage by 3000%, and enabled predictive staffing to improve patient care. - Worked with orchestration tools similar to Airflow and utilized cloud technologies (Microsoft Azure and AWS) for data infrastructure, achieving seamless cloud operations and reducing deployment times. - Leveraged strong Python and SQL expertise for ETL/ELT processes, developing scalable solutions with continuous improvement and adhering to best coding practices. Technologies: Python, SQL, JavaScript, Java, Azure ### Software Engineer - Data, AcuStream | R1 04/2013 - 09/2016 | Boulder, CO | https://www.r1rcm.com/ R1 RCM is an American healthcare revenue cycle management company servicing hospitals, health systems and physician groups across the United States. Headquartered in Chicago, Illinois, R1 RCM is publicly traded on the NASDAQ. - Built a custom invoicing system leveraging rule-based algorithms and machine learning, driving over $300M in annual revenue. - Migrated clients from SFTP to real-time APIs with Flask and Apache Kafka, reducing data delivery times by 50%. - Developed a custom invoicing system using Python and rule-based algorithms with ML components, generating $300M+ in annual revenue. Improved invoice accuracy by 35% and automated processes to reduce manual effort by 60%. - Partnered with executives on high-impact initiatives, driving a 200% increase in revenue and boosting client retention by 25%. Optimized operational strategies, improving profit margins from 41% to 78%. - Collaborated with CFOs and revenue cycle directors at leading healthcare systems to implement automated reconciliation processes. Cut reconciliation times by 50% and achieved 98% client satisfaction through scalable data solutions. - Built a financial reconciliation platform with Django and Celery, automating line-item invoicing and scheduling. Improved billing speed by 40% and streamlined operations despite complex client requirements. - Led infrastructure migration to AWS, moving PostgreSQL to RDS and compute workloads to EC2. Reduced query latency by 30% and cut infrastructure costs by 25%. Implemented IAM-based security, improving compliance audit performance by 20%. - Refactored 100+ code modules into optimized Python, leveraging multithreading, hashing, and compression techniques. Boosted system efficiency by 45%, resolving bottlenecks in data processing workflows. - Diagnosed and resolved issues in legacy Java applications, reducing downtime incidents by 15%. Enhanced UI responsiveness by 20% through performance optimizations in JavaScript. - Created a Django-powered cron job system with Celery for task orchestration and real-time tracking. Increased task completion rates by 30% through automated monitoring and recovery mechanisms. - Migrated clients from SFTP to real-time APIs with Flask and Apache Kafka, reducing data delivery times by 50%. Enhanced data accessibility, streamlining client operations and improving service quality. - Developed fault-tolerant systems with Spring Boot, implementing state consistency mechanisms such as transaction rollbacks. Reduced system fault impact by 35%, ensuring high availability during critical operations. - Supported pre-sales efforts and customer onboarding with on-site integrations, accelerating implementation timelines by 20% and enhancing customer satisfaction. - Technologies Used: Python, Java, JavaScript, SQL, Django, Flask, Celery, AWS (RDS, EC2, IAM), Apache Kafka, Spring Boot, PostgreSQL. Technologies: Python, SQL, JavaScript, Java, AWS ## Projects ### AI Billing System A proof-of-concept application for tracking and billing AI chat thread interactions with real-time cost metrics and analytics. This application is designed to solve the challenge of monitoring and billing for AI model usage in chat-based applications. As organizations increasingly deploy AI assistants, understanding the costs associated with these interactions becomes critical for business planning and cost management. Link: https://github.com/jckail/ai_chat_billing_app Technologies: Python, FastAPI, React, Pydantic, SQLAlchemy, Supabase, Docker ### Super Teacher AI assistant to help teachers manage their classroom and create personalized lesson plans for students. Made with Python, FastAPI, and Pydantic. Developed Super Teacher, an AI-powered web application that revolutionizes lesson planning for educators by providing personalized lesson plans tailored to individual student needs. Link: https://github.com/jckail/superteacher | Live: https://www.the-super-teacher.com/ Technologies: Python, FastAPI, React, Pydantic, SQLAlchemy, Supabase, Docker, GCP ### Join Group via QR Featured on TechCrunch - Won internal hackathon and created the ability for Facebook group admins to invite users to their groups by generating a QR Code. Used by millions daily. Link: https://techcrunch.com/2022/03/09/facebook-rolls-out-new-tools-for-group-admins-to-manage-their-communities-and-reduce-misinformation/ ### Jobbr Agentic AI job matching a resume to available jobs at tech companies. Scrapes 100,000 jobs in under 15 minutes, using Python, Langchain, FastAPI and OpenAI. Developed Jobbr, an AI-powered web application that transforms recruitment by matching job postings with candidates' skills and resumes for greater efficiency and accuracy. Leveraged GPT-4 and Anthropic’s Claude to parse job postings from raw HTML into structured JSON, aligning job descriptions with resumes using advanced NLP techniques. Integrated SQLModel and SQLAlchemy for efficient data management, with Alembic migrations to ensure smooth database updates and version control. Deployed the platform using Docker to maintain consistent environments and follow DevOps best practices. Implemented a FastAPI-based modular API with private endpoints, Supabase authentication, and AI-powered role parsing and URL scraping to promote code reusability and scalability. Used LangChain for document embedding and similarity analysis, enhancing job matching precision with advanced AI capabilities. Followed CI/CD workflows with comprehensive testing to ensure code reliability and seamless deployments. Prioritized security through token management, Supabase authentication, and password hashing with Python-JOSE and Passlib, alongside email validation to ensure strong user management. Employed Pandas, BeautifulSoup4, lxml, and NumPy for data processing and parsing job postings efficiently. Demonstrated proficiency across AI/ML, web development, data engineering, and cloud integration, solving real-world recruitment challenges with an AI-first approach. Tech Stack: Python 3.9+, FastAPI, SQLAlchemy, Pydantic, Supabase, Docker Link: https://github.com/jckail/Jobbr Technologies: Python, FastAPI, LangChain, OpenAI, Web Scraping ### Portfolio Website A fully custom-built portfolio site showcasing my skills and work. Made with FastAPI and React, hosted on GCP via CloudRun. This website is a custom-built portfolio site that I created to showcase my work. It is built using FastAPI and React, and is hosted on Google Cloud Platform via CloudRun. I chose React and TypeScript on the front end with a FastAPI backend so the whole stack stays typed end to end, and deployed it to Cloud Run behind Terraform-managed infrastructure. The site features a clean and modern design, with sections for my projects, skills, and experience. It also includes a contact form that allows visitors to get in touch with me. Overall, I am very happy with how the site turned out, and I think it does a great job of highlighting my work and skills. Link: https://github.com/jckail/portfolio Technologies: TypeScript, React, Vite, Python, FastAPI, Pydantic, Docker, GCP Cloud Run, Terraform, Supabase, Anthropic Claude ### Pointup.io An AI web app to manage 'All of your loyalty points in one place,' powered by a Selenium-WebDriver agent in AWS via elastic beanstalk, lambda, and s3. Created PointUp, a loyalty management tool designed to centralize and optimize points from credit cards, airlines, and hotel programs. Engineered a backend system using Python and Selenium for web scraping, overcoming anti-bot protections with NordVPN IP rotation and human behavior simulation. Deployed the system on AWS EC2 with secure data storage in S3, ensuring scalability and reliability. Integrated Auth0 for user authentication and encrypted sensitive data, ensuring secure handling of user credentials. Designed modular bot classes for easy integration of new loyalty programs, enhancing extensibility and user experience. Overcame complex anti-scraping challenges while maintaining ethical considerations around automation. Demonstrated expertise in Python development, cloud infrastructure, and secure web scraping through a real-world solution. Link: https://github.com/jckail/point_bot | Live: https://www.pointup.io/ Technologies: Python, TypeScript, Selenium, AWS Lambda, AWS S3, Elastic Beanstalk, Docker ### Algo Crypto Trading algorithm via scraped crypto, NASDAQ, and CPME rare minerals data. Made with Flask, Pandas, AWS, and SageMaker. Developed a Python-based data analysis tool to collect, process, and analyze real-time cryptocurrency data from APIs (CryptoCompare, CoinMarketCap, Alpha Vantage). Engineered a multithreaded data ingestion system to efficiently handle multiple data sources and ensure low-latency processing. Leveraged AWS services (S3, Glue, Athena) to create a scalable cloud-native architecture capable of handling large datasets. Applied stepwise regression and K-means clustering to detect market trends and correlations between traditional assets and cryptocurrencies. Demonstrated advanced data engineering skills by building modular components focused on data acquisition and ML-powered reporting. Highlighted problem-solving and financial technology expertise, navigating API integration challenges and delivering actionable insights into cryptocurrency trends. Link: https://github.com/jckail/crypto_trader Technologies: Python, Flask, Pandas, AWS, SageMaker ### goPilot Agentic AI developer assistant and cli tool to help debug, develop, and compile apps written in GO. Built using python and go leveraging OpenAI assistants api. Built GoPilot, a next-generation CLI tool to enhance development workflows and accelerate mastery of Go programming. Integrated GPT-4 for code insights and debugging, transforming the CLI into an intelligent development companion. Designed a modular system combining Go, Python, and Bash to manage file operations, context updates, and automated testing workflows. Implemented goci-lint and Go’s native testing framework to streamline linting and testing operations. Added web scraping capabilities to fetch the latest Go documentation, ensuring developers stay updated on best practices. Developed a customizable CLI interface with task toggles and output path definitions to adapt to diverse workflows. Demonstrated proficiency in Go, Bash scripting, and AI-enhanced development, showcasing innovative problem-solving and productivity enhancement. Link: https://github.com/jckail/goPilot Technologies: Python, Go, OpenAI API, CLI ### Data Playground An interactive data-engineering lab for pipeline DAGs, data models, SQL analytics, graphs, and vector similarity. Built a reproducible synthetic commerce pipeline with seeded event generation, deduplication, schema and lifecycle checks, and SQL analytics. Explore acquisition and retention scenarios, inspect quarantined records, and trace conversion, collected revenue, and cohort retention to the queries that compute them. The public React lab serves generated Python runs without a database; a bounded FastAPI service supports custom simulations when configured. A separate synthetic commerce dataset connects products, customers, and purchases in a relationship graph and explains cosine similarity over handcrafted product feature vectors. An engineering workbench exposes executable dependency DAGs, retry and failure traces, model grains and contracts, and architecture decisions grounded in the implementation. The original PostgreSQL and Streamlit experiment remains in the source repository. Link: https://github.com/jckail/data_playground | Live: https://jckail.com/dataplayground Technologies: Python, SQL, SQLite, FastAPI, Pydantic, React, TypeScript, Docker ## Skills ### Data Engineering - **Airbyte** (Data Integration): Open-source data integration platform with a large connector catalog for syncing data from APIs, databases and SaaS tools into warehouses and lakes. Commonly used to stand up ELT ingestion without hand-written extract code. - **Airflow** (Workflow Orchestration): Python-based workflow orchestrator that defines pipelines as DAGs with scheduling, retries and monitoring. A standard choice for batch data pipelines, ML training schedules and RAG ingestion jobs. - **Kafka** (Stream Processing): Distributed commit-log platform for high-throughput event streaming. The usual backbone for real-time pipelines, change-data capture and event-driven architectures. - **Pulsar** (Stream Processing): Cloud-native messaging and streaming platform with multi-tenancy and tiered storage. An alternative to Kafka for pub/sub plus queue workloads. - **Cloud Composer** (Workflow Orchestration): Google Cloud's managed Apache Airflow service. Used to run DAG-based pipelines without operating the scheduler and workers yourself. - **Cloud Data Fusion** (Data Integration): Managed, visual ETL/ELT service on Google Cloud built on the open-source CDAP project. Used to assemble batch data integration pipelines with prebuilt connectors. - **Dataprep** (Data Preparation): Visual data-preparation service on Google Cloud for profiling, cleaning and transforming data without code. Used for ad-hoc preparation ahead of analytics or ML. - **Dataproc** (Data Processing): Managed Spark and Hadoop service on Google Cloud with fast cluster start-up. Used for batch ETL and ML data preparation at scale. - **DBT** (Data Transformation): SQL-first transformation framework that adds version control, testing and documentation to warehouse models. Standard tool for analytics engineering and data-quality checks. - **Flink** (Stream Processing): Stateful stream-processing engine with event-time semantics and exactly-once guarantees. Used for real-time pipelines, enrichment and streaming analytics. - **Kinesis Firehose** (Stream Processing): Managed AWS service that batches and delivers streaming data to S3, Redshift, OpenSearch and other destinations. Used for low-ops stream ingestion. - **PubSub** (Messaging): Google Cloud's managed publish/subscribe messaging service for asynchronous, event-driven systems. Used to decouple producers and consumers at scale. ### Development Tools - **Airtable** (Collaboration Platform): Cloud spreadsheet-database hybrid with views, forms and automations. Often used for lightweight operational tracking and internal workflows before data moves to a warehouse. - **Figma** (UI/UX Design): Browser-based interface design and prototyping tool. Used for UI mockups, design systems and design hand-off to engineers. - **Linux** (Operating System): Open-source operating system kernel and ecosystem that runs most servers and containers. Core to day-to-day infrastructure and production debugging. - **Playwright** (Testing Framework): Browser automation and end-to-end test framework for Chromium, Firefox and WebKit. Used for reliable UI smoke and regression tests. - **Selenium** (Testing Framework): Long-standing browser automation framework driven by WebDriver. Used for end-to-end UI testing and web scraping. - **Retool** (Low-Code Platform): Low-code platform for assembling internal tools and admin panels on top of databases and APIs. Used for operational dashboards and back-office workflows. ### Artificial Intelligence - **Amazon Bedrock** (Cloud AI Services): Managed AWS service that exposes foundation models from several providers behind one API, with guardrails and agent tooling. Typically used to add LLM features to AWS-hosted applications without running model infrastructure. - **Chroma** (Vector Database): Open-source embedding database for storing and querying vectors with metadata. Popular for prototyping retrieval-augmented generation (RAG) applications. - **Claude AI** (Language Models): Family of large language models and assistant from Anthropic, strong at long-context reasoning, coding and tool use. Used to build assistants and agents through the API. - **Google Gemini** (Language Models): Google's family of multimodal foundation models for text, code, images and more, available through Vertex AI and the Gemini API. Used for assistants, tool-calling agents and long-context tasks. - **Hugging Face** (Machine Learning Platform): Hub and libraries (Transformers, Datasets) for sharing, fine-tuning and serving open models. Central to working with open-weight NLP and LLM models. - **Keras** (Deep Learning Framework): High-level deep-learning API that runs on top of TensorFlow and other backends. Used to define and train neural networks quickly. - **LangChain** (LLM Framework): Framework for composing LLM applications from prompts, retrievers, tools and agents. Commonly used for RAG pipelines and tool-using agents. - **Mistral AI** (Language Models): Developer of efficient language models, with several open-weight releases. Used where smaller or self-hostable models are preferred. - **OpenAI** (AI Services): AI lab and API provider of GPT models, embeddings and tool-calling. Used to build chat assistants, RAG systems and agents. - **PyTorch** (Deep Learning Framework): Open-source deep-learning framework with dynamic graphs and a research-to-production path. Used to train, fine-tune and serve neural models, including classifiers and LLMs. - **TensorFlow** (Deep Learning Framework): Google's end-to-end machine-learning platform for training and serving models. Widely used for production deep learning and mobile or edge inference. - **Vertex AI** (ML Platform): Google Cloud's managed ML platform for training, tuning, deploying and evaluating models, including Gemini. Used to run ML and LLM workloads on GCP. - **Agent Platforms** (Agentic Systems): Shared infrastructure for building, running and governing LLM-powered agents: model access, tool and permission management, memory, orchestration, observability and evaluation. A platform lets many teams ship agents consistently instead of rebuilding the same plumbing. - **Agent Harnesses** (Agentic Systems): The runtime scaffolding around a model that turns it into an agent: the loop that calls tools, manages context and state, enforces guardrails and permissions, and records traces. A good harness makes agent behavior repeatable, testable and safe to operate. - **Agent Evaluation** (Evaluation): Testing agents end to end: scored task suites, tool-call and trajectory checks, tracing and replay, and regression gates before release. Because agents act over many steps, evaluation looks at the whole trajectory rather than a single answer. - **LLM Evaluation** (Evaluation): Measuring language-model quality with curated datasets, rubric or model-graded scoring and human review, tracked over time to catch regressions when prompts, models or retrieval change. - **AI Classifiers** (Machine Learning Systems): Model-based systems that assign labels or scores to content or events, from classical ML to transformer models. Production use depends on labeled data, feature and training pipelines, threshold calibration, evaluation and monitoring. - **Model Governance** (Responsible AI): Controls that keep statistical and ML models trustworthy in production: versioning, approval and audit trails, monitoring for drift and bias, documented decision logic, and clear ownership. Especially important for high-stakes domains such as identity and fraud. - **RAG** (LLM Frameworks): Retrieval-Augmented Generation: retrieving relevant documents, typically through embeddings and a vector store, and giving them to a language model as context. Grounds answers in current, private data and reduces hallucination. ### Cloud Computing - **Azure** (Cloud Platform): Microsoft's cloud platform covering compute, storage, data and AI services. Used for enterprise workloads, including managed data warehousing and hosted model APIs. - **AWS** (Cloud Platform): Amazon's cloud platform spanning compute, storage, databases, analytics and ML services. A common home for data lakes (S3, Athena), streaming and model-serving workloads. - **Google Cloud** (Cloud Platform): Google's cloud platform with data, analytics and AI services such as BigQuery, Pub/Sub and Vertex AI. Used to build data platforms and model workloads. ### Databases - **Cassandra** (NoSQL Databases): Distributed wide-column NoSQL database built for high write throughput and horizontal scale across data centers. Used for time-ordered, high-volume event and profile data. - **DuckDB** (Analytical Database): In-process analytical SQL database that reads Parquet, CSV and more directly. Useful for fast local analysis and as an embedded engine in data tools. - **DynamoDB** (NoSQL Database): AWS's fully managed key-value and document database with single-digit-millisecond latency at scale. Used for session, metadata and high-throughput lookup data. - **Elasticsearch** (Search Engine): Distributed search and analytics engine built on Lucene, supporting full-text, filtering and vector search. Used for log analytics and search features. - **Firestore** (NoSQL Database): Google Cloud's serverless document database with real-time sync. Used for app state and metadata that needs simple scaling. - **MongoDB** (Document Store): Document database that stores JSON-like documents with flexible schemas. Used for evolving, nested application data. - **MySQL** (Relational Database): Widely deployed open-source relational database. Used for transactional application data and as a reliable SQL back end. - **Neo4j** (Graph Database): Native graph database queried with Cypher, optimised for relationship-heavy data. Used for entity resolution, recommendations and knowledge graphs. - **Pinecone** (Vector Database): Managed vector database for low-latency similarity search over embeddings. Used for semantic search and RAG retrieval. - **PostgreSQL** (Relational Database): Open-source relational database with strong SQL, transactions, JSON and extension support. A dependable default for application and analytical data. - **Qdrant** (Vector Database): Qdrant is an open-source vector database and similarity search engine with payload filtering. Used for semantic search and RAG retrieval. - **Redis** (In-Memory Store): In-memory data store used as a cache, queue, rate limiter and lightweight database. Used to cut latency and absorb load in front of slower systems. - **RocksDB** (Embedded Database): Embedded key-value store based on LSM trees, tuned for fast writes and SSDs. Underlies stateful systems such as Flink state and Kafka Streams. - **Supabase** (Database Platform): Open-source back-end platform built on PostgreSQL, offering auth, storage and real-time APIs. Used to ship application back ends quickly. - **TimescaleDB** (Time Series Database): PostgreSQL extension for time-series data with automatic partitioning and compression. Used for metrics, telemetry and event data. ### Programming Languages - **C++** (Systems Programming): Compiled systems language that offers fine-grained control over memory and performance. Used for performance-critical engines, and underlies many ML and database internals. - **Go** (Systems Programming): Statically typed, compiled language with built-in concurrency, designed for simple and reliable services. Used for APIs, streaming consumers and infrastructure tooling. - **Java** (Enterprise Development): Mature object-oriented language on the JVM, widely used in enterprise back ends and big-data systems such as Kafka, Hadoop and Spark. - **JavaScript** (Web Development): The language of the web, running in browsers and on servers via Node.js. Used for interactive front ends and full-stack applications. - **PHP** (Web Development): Server-side scripting language that powers a large part of the web. Used for web applications and CMS-driven sites. - **Python** (General Purpose): General-purpose language with a rich ecosystem for data, ML and automation. The primary language for pipelines, model code and agent frameworks. - **Rust** (Systems Programming): Memory-safe systems language with performance comparable to C and C++. Used for high-performance services and data tooling. - **Scala** (JVM Languages): JVM language blending object-oriented and functional programming. Heavily used for Spark jobs and large-scale data processing. - **SQL** (Query Language): Standard language for querying and shaping relational data. Foundation of analytics, data modeling and pipeline logic. - **TypeScript** (Web Development): Typed superset of JavaScript that catches errors at compile time. Used for maintainable front ends and Node services. ### Big Data - **Hadoop** (Distributed Computing): Original open-source stack for distributed storage (HDFS) and batch processing across commodity clusters. Foundation of many data lakes that later moved to Spark and open table formats. - **Iceberg** (Data Lake Technology): Open table format that adds schema evolution, snapshots and ACID semantics to large analytic datasets on object storage. Lets Spark, Trino and Flink share the same lake tables. - **Spark** (Data Processing): Distributed analytics engine for large-scale batch and streaming processing with SQL, DataFrame and ML APIs. Widely used for ETL, feature engineering and training-data preparation. - **BigQuery** (Data Warehouse): Serverless columnar data warehouse on Google Cloud that runs SQL over very large datasets. Used for analytics, feature tables and increasingly ML directly in SQL. - **Databricks** (Data Platform): Lakehouse platform built around Spark with managed notebooks, job orchestration and MLflow-style ML tooling. Used for large-scale data engineering and model development. - **Presto** (Query Engine): Distributed SQL query engine for fast interactive analytics across many data sources. Originated at Facebook and is the ancestor of Trino. - **Snowflake** (Data Warehouse): Cloud data warehouse that separates storage and compute with elastic scaling. Used for governed analytics and data sharing. - **Trino** (Query Engine): Distributed SQL engine that queries many data sources in place, forked from Presto. Used for federated analytics over lakes and warehouses. ### Web Development - **CSS3** (Frontend Styling): Style language for the layout, typography and theming of web pages. Underpins responsive design, dark mode via custom properties, and accessible UI. - **Django** (Backend Framework): Batteries-included Python web framework with ORM, admin and auth. Used for data-backed web applications and internal tools. - **FastAPI** (API Framework): Modern Python web framework built on type hints and Pydantic, with automatic OpenAPI docs and async support. A common choice for model-serving and agent back ends. - **Flask** (Backend Framework): Lightweight Python micro-framework for web services. Often used for small APIs and for wrapping models behind HTTP endpoints. - **Gin** (Backend Framework): Fast HTTP web framework for Go with minimal overhead. Used to build lean, high-throughput APIs and services. - **GraphQL** (API Technology): Query language and runtime that lets clients request exactly the fields they need from a typed schema. Used to unify data access across services. - **HTML5** (Frontend Markup): Current HTML standard for structuring web content with semantic elements and built-in media and form support. The base layer for accessible web interfaces. - **Material UI** (UI Framework): React component library implementing Google's Material Design. Used for accessible, themeable UI building blocks. - **Node.js** (Runtime Environment): JavaScript runtime built on V8 for event-driven network services and tooling. Used for APIs, build pipelines and server-side rendering. - **React** (Frontend Framework): Component-based JavaScript library for building user interfaces. Used for interactive applications such as dashboards and chat UIs. - **Redux** (State Management): Predictable state container for JavaScript applications, commonly used with React. Used to manage complex client-side state. - **SQLAlchemy** (ORM): Python SQL toolkit and ORM with a Core expression layer. Used to talk to relational databases from services and pipelines. - **Svelte** (Frontend Framework): Compiler-based UI framework that ships minimal runtime JavaScript. Used for fast, lightweight web front ends. - **Tailwind CSS** (CSS Framework): Utility-first CSS framework for composing designs directly in markup. Used for consistent, rapid UI styling. - **Vite** (Build Tools): Fast front-end build tool with instant dev server and optimised production bundling. Used to build modern single-page apps. - **WebAssembly** (Web Standards): Portable binary instruction format that runs near-native code in browsers and other runtimes. Used to bring compute-heavy code, from Rust or C++, to the web. ### DevOps - **Datadog** (Monitoring): SaaS observability platform for metrics, traces, logs and alerting. Used to monitor services and data pipelines and to drive on-call alerting. - **Docker** (Containerization): Container tooling that packages an application with its dependencies into a portable image. Used to ship reproducible services, pipelines and model-serving workloads. - **Git** (Version Control): Distributed version control system for tracking changes and collaborating on code. Foundation of review, branching and release workflows. - **GitHub Actions** (CI/CD): CI/CD and automation built into GitHub, defined as workflow files in the repository. Used to test, scan and deploy services and to keep builds reproducible. - **Grafana** (Monitoring & Visualization): Open-source dashboarding and alerting layer over metrics, logs and traces. Commonly paired with Prometheus for service and pipeline monitoring. - **Jaeger** (Distributed Tracing): Open-source distributed tracing backend for following a request across services. Used to find latency and error sources in microservice systems. - **Kubernetes** (Container Orchestration): Container orchestrator that schedules, scales and heals workloads across a cluster. The common substrate for serving services, jobs and ML or GPU workloads. - **NGINX** (Web Server): High-performance web server and reverse proxy with load balancing and caching. Commonly the front door for web apps and APIs. - **OpenTelemetry** (Observability): Vendor-neutral standard and SDKs for emitting traces, metrics and logs. Used to instrument services once and route telemetry to any backend. - **Prometheus** (Monitoring): Time-series monitoring system with a pull model, PromQL and alerting rules. A standard for service and Kubernetes metrics. - **Splunk** (Log Management): Platform for searching, analysing and alerting on machine data and logs. Used for operational monitoring, security and incident investigation. - **Terraform** (Infrastructure as Code): Infrastructure-as-code tool that declares cloud resources in configuration and plans changes before applying them. Used for reproducible, reviewable infrastructure. ### Data Science - **Jupyter** (Interactive Computing): Interactive notebooks that mix code, results and narrative. A standard workspace for data exploration, model prototyping and sharing analysis. - **Looker** (Business Intelligence): Business-intelligence platform with a governed semantic layer (LookML). Used to give teams consistent metrics and self-serve dashboards. - **NumPy** (Scientific Computing): Core Python library for fast n-dimensional array computing. Foundation for pandas, scikit-learn and most numerical and ML code. - **Pandas** (Data Analysis): Python library for tabular data analysis with DataFrames. Used for data cleaning, exploration and feature preparation. - **Scikit-learn** (Machine Learning): Python library of classical machine-learning algorithms with a consistent API. Used for classification, regression, clustering and evaluation baselines. - **Streamlit** (Data Apps): Python framework for building interactive data and ML apps with minimal front-end code. Used for quick demos, internal dashboards and model explorers. - **Tableau** (Business Intelligence): Visual analytics platform for building interactive dashboards. Used to explore data and communicate metrics to business teams. ## Links - GitHub: https://github.com/jckail - LinkedIn: https://www.linkedin.com/in/jckail/ - Website: https://www.jckail.com/ - Resume (PDF): https://www.jckail.com/api/resume - Resume (JSON Resume): https://www.jckail.com/resume.json - Location: Denver, CO, United States