| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.
You must be logged in to block users.
Contact GitHub support about this user’s behavior. Learn more about reporting abuse.
Report abuseI build data science and machine learning systems that move cleanly from raw data to defensible insight, with an emphasis on well-motivated problems, reproducible pipelines, and model interpretability.
My work sits at the intersection of applied ML, scientific computing, and software engineering, often using public or operational data to prototype end-to-end analyses that could realistically run in production.
Distributed-ML
A distributed, dataset-agnostic CT preprocessing pipeline using Dask, designed for large clinical imaging datasets and downstream ML workflows.
publicdata_ca
A reusable data acquisition and normalization framework for Canadian public datasets (StatCan, CMHC, CIHI), supporting rapid ML case studies such as housing affordability indices and hospital utilization analysis.
Applied ML Case Studies
Short, tightly scoped projects demonstrating:
YesChef GPT
An AI-powered system that structures generative outputs into machine-readable components (ingredients, preparation steps, pickup notes), emphasizing controllability and downstream usability over novelty.
Anomaly detection in healthcare operations
Early detection of unusual demand or utilization patterns using unsupervised and semi-supervised methods.
Public-sector ML pipelines
Designing reusable ingestion and feature pipelines that make public data viable for rapid experimentation.
Evaluation without labels
Practical techniques for validating unsupervised models when ground truth is incomplete or unavailable.
Bridging notebooks to systems
Turning exploratory analyses into maintainable, testable services without losing scientific intent.
I approach data science as an engineering discipline:
start with a clear question, respect the data’s limitations, and build models that can be explained, tested, and trusted.
My goal is to work on problems where statistical thinking, ML techniques, and real-world constraints all matter — especially in healthcare, infrastructure, and public data contexts.
Housing Affordability Stress Index
Jupyter Notebook
Distributed-ML: CT Preprocessing Pipeline with Dask This repository implements a distributed, dataset-agnostic CT preprocessing pipeline designed for large clinical imaging datasets such as NLST, C…
Python
An AI-powered resume builder based on Martin Yate’s Knock ’Em Dead formula. It finds job ads, extracts key skills, and helps craft tailored, achievement-driven resumes optimized for each role, with…
Python
EVXchange is a web app that lets electric vehicle (EV) owners find and book nearby charging stations hosted by individuals or businesses. It works like AirBnB, but for EV chargers — people can rent…
Python 1
| Back | FazBrowse Home | New Git URL |