Clinical Readmission Risk Modeling
A leakage-safe 30-day readmission pipeline with patient-disjoint cohorts, calibration, missingness analysis, and drift checks.
Held-out test: 0.661 ROC-AUC, 0.175 PR-AUC, and 0.077 Brier score on 10,822 encounters.Yogesh Kuchimanchi · Data Scientist
I’m Yogesh, a data science graduate student at RIT. Collecting parts got me interested in the numbers behind a decision. These days, that curiosity takes me into machine learning, language, and how things change over time.
A little about me
I like building PCs. Somewhere between collecting parts and comparing them, I started enjoying the comparisons as much as the build itself. Why this component? What am I getting for the difference in price? Will it make a difference for what I want to do?
That is how I got interested in data. I’m now studying data science at Rochester Institute of Technology, and those questions have followed me into my projects.
Choosing parts means making trade-offs. A bigger number on a spec sheet only tells you so much; the workload and the rest of the build matter too. That is a question I now bring to models: did the comparison give each one a fair test?
Collecting parts also means collecting information: specs, reviews, recommendations. Making sense of all that is part of what drew me to data. With text, the same problem gets much bigger.
A PC build is a choice made at a particular moment. Workloads change, and what was enough before may not be enough later. I’m interested in that with models too: what happens when the data they meet changes?
At RIT, my research takes that interest in change into a very different setting: public conversations about women’s safety in India. The question is how those conversations change after an incident, and what remains months or years later.
One question, at different moments
These are the three time windows in our women’s safety research. Choose a window to see the question it helps us ask.
The crossroads
Research gives me room to investigate a question properly. Engineering lets me turn what I learn into something someone can use. As I finish my master’s, I’m looking for a data science role where I can keep asking questions and building things.
Featured projects
Four projects to explore. There’s a small PC build below if you want to try it; each project is also linked directly underneath.
Explore all GitHub projectsA nod to where it started
A small interactive build, with each part representing a step in an ML project.
Select a component, then fit it into the matching slot.
Installed applications
A leakage-safe 30-day readmission pipeline with patient-disjoint cohorts, calibration, missingness analysis, and drift checks.
Held-out test: 0.661 ROC-AUC, 0.175 PR-AUC, and 0.077 Brier score on 10,822 encounters.A four-class AG News system with a DistilBERT training pipeline, FastAPI service, React dashboard, and explicit inference modes.
Held-out evaluation: 0.870 accuracy, 0.869 macro-F1, and 0.827 MCC on 12,000 articles.A deterministic simulator for performance decay, class-prior shift, PSI, approximate KS, and configurable monitoring alerts.
Verified mixed-shift run: ROC-AUC fell from 0.758 to 0.372; max PSI reached 0.608; three alerts fired.An interactive research dashboard covering 351,501 Reddit and YouTube comments across 16 women-safety cases in India.
Paper accepted at ASONAM 2026. The dashboard separates submitted-paper findings from later model audits.The full collection
Explore my public repositories, from machine learning systems to experiments and tools. New public repositories appear automatically.
46 of 46 repositories · Most recently updated first · Refreshing from GitHub…
Deployable FastAPI RAG evaluation API with source citations, vector retrieval, Docker, and reproducible metrics.
Inference cost planning, PyTorch classification, document retrieval, and LLM evaluation tools.
Full-stack news aggregator with DistilBERT summarization, React dashboard, CI/CD pipeline, and Docker deployment
Explore the source code and project files on GitHub.
Explore the source code and project files on GitHub.
Product analytics and A/B testing project with Python, SQL, Streamlit, and reproducible metrics.
Explore the source code and project files on GitHub.
Bounded self-healing data pipeline with LangGraph, strict contracts, isolated repair, durable recovery, and an evidence-based technical report.
Explore the source code and project files on GitHub.
Explore the source code and project files on GitHub.
A modern, interactive portfolio website built with Next.js, TypeScript, and Tailwind CSS
FixMyData is a smart AI-powered application designed to automatically detect and explain data quality issues in CSV files including missing values, outliers, duplicate entries, type mismatches, and more.
Explore the source code and project files on GitHub.
Retrieval-augmented generation pipeline with sentence-transformer embeddings, ChromaDB vector search, and TinyLlama synthesis
FastAPI code review server powered by TinyLlama-1.1B with streaming SSE, severity classification, and KV-cache optimization
Image-to-text generation using ViT encoder and GPT-2 decoder with beam search, served via FastAPI
Real-time sentiment analysis API with DistilBERT, batch processing, and trend visualization via FastAPI and Streamlit
Multi-agent LLM framework with planner, domain experts, and synthesizer using chain-of-thought reasoning
A machine learning pipeline to predict credit card default, exposed as a RESTful API using FastAPI.
AI-powered restaurant analytics dashboard with natural language SQL queries. Built with Streamlit, SQLite, and LLM integration for intuitive data exploration and insights.
Real-time model monitoring dashboard with KS test, PSI, and Jensen-Shannon divergence for data drift detection
Neural style transfer with VGG19 and procedural generative art including fractals, flow fields, and wave interference
NER-driven knowledge graph construction with interactive PyVis visualization and community detection
Multi-model time series benchmarking with ARIMA, Holt-Winters, and automated feature engineering
Ensemble anomaly detection with Isolation Forest, LOF, DBSCAN, and synthetic data augmentation for imbalanced datasets
Explore the source code and project files on GitHub.
Professional AI-powered symptom analysis tool using Ollama and Streamlit. Provides structured medical insights with comprehensive safety disclaimers for educational purposes.
Real-time dashboard to track public sentiment around stock tickers and compare it with actual price trends using NLP and Yahoo Finance data.
Advanced Metal Surface Defect Detection System using PyTorch with CNN, Attention Mechanisms, Ensemble Learning, SMOTE, and Cross-Validation
Explore the source code and project files on GitHub.
Explore the source code and project files on GitHub.
🎬 Advanced Deep Learning for IMDB Movie Review Sentiment Analysis - Achieving 90%+ accuracy with LSTM, BiLSTM, CNN, and Hybrid models
Explore the source code and project files on GitHub.
🎨 An AI-powered creative platform that generates art, analyzes emotions, and creates narratives through an intuitive web interface. Built with React, Flask, and multiple AI APIs.
Explore the source code and project files on GitHub.
🧠🎨 AI-Powered Emotional Art Generation Platform - Transform emotions into stunning neuromorphic art through advanced multimodal AI analysis with contextual memory and narrative generation
An immersive AI-powered historical storytelling web application that transports users through different eras with multimodal narratives, user authentication, and interactive features.
AI-Powered Slide Deck Generator - Create professional presentations from text using Google Gemini AI with multiple styles, charts, and speaker notes
🎭 Echo-Muse: AI-Powered Therapeutic Storytelling Companion - Personalized healing through AI-generated stories and ambient soundscapes
AI-powered platform that generates therapeutic soundscapes from plant bioacoustic data
An AI-powered exploration of emotional connections through historical letters, visualized as living fungal networks
Python Script to check file duplicacies
A Flask web application for image recognition using Ollama and a multimodal model.
Machine Learning project for ml-iris-data-analysis
Machine Learning project for ml-digits-exploration
Chrun Based Prediction
Experience
Rochester Institute of Technology
Rochester Institute of Technology
Rochester Institute of Technology
Education
Rochester Institute of Technology · Rochester, NY
Machine Learning, Deep Learning, Cloud Computing, Big Data AnalyticsVellore Institute of Technology · India
New Shores International College · India
Skills
Grouped by the work they support, not by logo count.
Python · SQL · R · Java
PyTorch · scikit-learn · CatBoost · XGBoost · Transformers · LoRA · DistilBERT · FT-Transformer
A/B testing · Mann–Whitney U · Chi-square · G-test · ROC-AUC · PR-AUC · Calibration · PSI / KS / JS
Pandas · NumPy · SciPy · Spark · PostgreSQL · SQLite · Window functions
FastAPI · Streamlit · React · Docker · GitHub Actions · Automated testing · Git · Excel / VBA
Contact
I’m open to data science and ML engineering opportunities, research collaboration, and useful technical conversations.