SSK.
Open to opportunities · USA
AI & Machine Learning Engineer / Data Scientist

Sai Sandeep Kethiboina

Five-plus years designing, building, and shipping production AI/ML, deep learning, and generative-AI systems across banking, healthcare, and GPU-accelerated enterprise environments — from raw data pipelines to LLM-powered assistants serving millions of records.

05+
Years in production ML
03
Enterprise domains
150M+
Records processed
90%+
LLM response relevance
Scroll · 00 → 06
01 Profile
Portrait of Sai Sandeep Kethiboina

AI/ML engineer who turns messy enterprise data into predictive systems and LLM products that move real business metrics.

I design end-to-end machine learning and generative-AI solutions — predictive analytics, deep learning, NLP, time-series forecasting, fraud detection, and RAG-based assistants — and carry them all the way to production with MLOps/LLMOps discipline. My work spans banking, healthcare, and GPU-driven enterprise systems, where reliability, governance, and explainability are non-negotiable. I care about the unglamorous parts: clean pipelines, monitored models, and outcomes you can measure.

DisciplineAI / ML Engineering
Experience5+ years in production ML
CurrentlyAI/ML Engineer · Capital One
EducationM.S. Computers & Information Science
DomainsBanking · Healthcare · GPU Computing
BasedTexas, USA
02 Experience
Jan 2025 — Present
McLean, VA, USA
Banking

AI / Machine Learning Engineer

Capital One
  • Improved predictive accuracy +30% across credit-risk, fraud, loan-default, and customer-segmentation models built with PyTorch, XGBoost, and LightGBM.
  • Reduced data-processing time −45% by engineering scalable ETL and real-time feature pipelines on Spark, Airflow, and Snowflake for high-volume transaction data.
  • Reached 90%+ response accuracy with RAG banking assistants over FAISS/Pinecone, using LangChain and the OpenAI API for governed policy and risk-framework retrieval.
  • Cut manual effort −40% by automating document summarization, dispute resolution, and compliance reporting with LoRA fine-tuned GenAI workflows.
  • Accelerated deployment +60% through enterprise MLOps/LLMOps on MLflow and Kubeflow — model registry, drift detection, and monitoring dashboards for compliance stakeholders.
Aug 2023 — Dec 2024
Louisville, KY, USA
Healthcare

Machine Learning Engineer

Humana
  • Lifted healthcare data availability +40% by building end-to-end AI/ML pipelines across 50M+ patient records with Spark, Airflow, and Databricks on PHI-ready FHIR/HL7 datasets.
  • Raised model accuracy +25% with deep-learning models for patient risk stratification, readmission prediction, and claims-fraud detection in TensorFlow and PyTorch.
  • Hit 90%+ response relevance on LLM clinical assistants using RAG and vector search over FAISS/Pinecone for secure guideline and patient-care retrieval.
  • Cut manual documentation effort −35% by automating clinical coding with LangChain and LoRA-fine-tuned models, plus +45% NLP/OCR processing on EHRs and physician notes.
  • Slashed deployment time −60% at 99.9% availability through HIPAA-aligned MLOps on MLflow and Kubeflow with responsible-AI validation and a CI/CD model registry.
May 2020 — Nov 2022
Hyderabad, India
GPU / Deep Learning

Data Scientist

Nvidia
  • Accelerated experimentation throughput by building GPU-accelerated data-science workflows with CUDA, RAPIDS, and PyTorch for scalable model analysis.
  • Improved model-serving reliability with TensorRT inference benchmarks across ONNX and Kubernetes for production-grade GPU performance comparisons.
  • Processed 100M+ structured and unstructured records through scalable ETL on Spark, Hadoop, Hive, and Snowflake for advanced analytics and AI.
  • Built anomaly-detection, time-series-forecasting, clustering, and recommendation models, with real-time streaming on Kafka and Spark Streaming.
  • Strengthened defect analysis and product insights via computer-vision models (OpenCV, TensorFlow) and telemetry analytics surfaced in Tableau and Power BI.
03 Selected Work

Open-source projects built end-to-end — from data and modeling to serving. Each links to its repository on GitHub.

LLM/01

Local RAG Chatbot

Fully local retrieval-augmented chatbot: ingests and embeds Markdown docs into a Chroma vector store and answers via a llama.cpp-served LLM. FastAPI backend with streaming chat, document upload, conversation memory, and incremental re-indexing; React + TypeScript frontend.

FastAPIllama.cppChromaDBReact
View repository ↗
Generative AI/02

Synthetic Tabular Data with GANs

TGAN/CTGAN generation of synthetic healthcare records, evaluated across statistical similarity (PCA, autoencoders, clustering), a custom privacy-at-risk metric, and downstream ML utility on length-of-stay and mortality prediction.

CTGANTGANTensorFlowscikit-learn
View repository ↗
MLOps/03

End-to-End MLOps on GCP

Reference pipeline: Prefect ETL of SF 311 data into BigQuery, dbt transforms, MLflow experiment tracking, and a scikit-learn model served via FastAPI on Cloud Run — provisioned with Terraform and shipped through GitHub Actions.

PrefectBigQuerydbtTerraformMLflow
View repository ↗
LLM/04

LLaMA From Scratch · 2.3M params

A 2.3M-parameter LLaMA-style language model implemented from scratch in PyTorch — RMSNorm, rotary positional embeddings (RoPE), and SwiGLU — trained on character-level TinyShakespeare.

PyTorchLLaMARoPETransformers
View repository ↗
Time Series/05

Time-Series Forecasting Benchmark

End-to-end forecasting on the Beijing PM2.5 dataset benchmarking ~25 approaches — from ARIMA/SARIMAX and Holt-Winters to XGBoost/LightGBM and LSTM/DeepAR/Prophet — scored on MAE, RMSE, MAPE, and R².

TensorFlowXGBoostLightGBMProphet
View repository ↗
Computer Vision/06

Computer Vision Collection

A set of CV notebooks: medical image classification (cataract, pneumonia, eye disease), traffic-sign and emotion recognition, driver-drowsiness detection, and OpenCV / YOLOv3 detection demos.

TensorFlow/KerasCNNOpenCVYOLOv3
View repository ↗
Recommenders/07

Multi-Modal E-commerce Recommender

Fashion recommender exploring four strategies — collaborative, content-based, hybrid, and a multi-modal PyTorch model fusing user/product embeddings with ResNet50 image and SentenceTransformer text features over a GCN layer.

PyTorchFlaskResNet50PyG
View repository ↗
Data Engineering/08

Airflow ETL → Snowflake (SCD2)

An Apache Airflow DAG extracting HR and salary data from PostgreSQL to S3, diffing against the warehouse, and loading into Snowflake with Slowly Changing Dimension Type 2 to track salary history over time.

AirflowPostgreSQLSnowflakeAWS S3
View repository ↗
04 Stack & Capabilities

Languages /01

PythonSQLRScalaJavaC++Bash

AI / ML /02

Deep LearningNLPComputer VisionReinforcement LearningTime-SeriesFeature Eng.

GenAI & LLMs /03

OpenAIHugging FaceLangChainLangGraphLlamaIndexRAGLoRA / PEFTAgentic AI

Frameworks /04

TensorFlowPyTorchScikit-learnKerasXGBoostLightGBMCatBoost

Data & Vector /05

Apache SparkKafkaHadoopAirflowSnowflakeFAISSPineconeChromaDB

Cloud & MLOps /06

AWS SageMakerVertex AIAzure MLDatabricksMLflowKubeflowDockerKubernetesCI/CD
05 Impact, by the numbers
150M+
Records processed in production pipelines
Humana + Nvidia
60%
Model deployment time via enterprise MLOps
Humana
+25%
Model accuracy on patient-risk & fraud detection
Humana
99.9%
Availability on production AI serving
Humana
45%
Data-processing time on banking pipelines
Capital One
+35%
Operational efficiency from NLP document intelligence
Capital One
06 Contact

Let's build something
measurable.