FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

data-splitting · GitHub Topics · GitHub

#

data-splitting

Here are 33 public repositories matching this topic...

MoleculeNet benchmark dataset & MolMapNet dataset

  • Updated Mar 29, 2022
  • HTML

A Python toolkit for file processing, text cleaning and data splitting. 文件处理,文本清洗和数据划分的python工具包。

  • Updated Oct 18, 2022
  • Python

An Exploratory Toolkit for Recommender Systems Datasets and Splits

  • Updated Mar 14, 2026
  • Jupyter Notebook

Data-Splitter is a Python script designed to split a large CSV file containing data into three different formats: JSON, a database table, and another CSV file. The script ensures a random distribution of data across the three output formats based on custom-defined ratios.

  • Updated Jul 24, 2023
  • Jupyter Notebook

Automated leakage-aware splitting of a given PPI dataset into train, validation, and test set

  • Updated Aug 18, 2026
  • Python

Code used for the analysis described in "Towards mobile music emotion recognition with cEEGrid", studying MER with a range of models, feature extraction, and data splitting techniques. Set up for the DEAP and DAAMEE datasets.

  • Updated Feb 17, 2026
  • Python

splitting image dataset into train, val, test sets

  • Updated Apr 23, 2023
  • Python

Reproducible archetype-based train/test splitting and benchmarking for ML data sets

  • Updated Feb 16, 2026
  • Python

Analyzed customer churn using transaction data. Built ML model to predict lapses. Dataset includes customer status, collection/redemption info, and program tenure. Delivered business presentation outlining modeling approach, findings, and churn reduction strategies.

  • Updated Apr 18, 2024

This project focuses on cleaning and analyzing a loan application dataset to gain insights into the factors influencing loan defaults. Through systematic data cleaning, visualization, and merging with previous application data, it provides a robust foundation for further predictive modeling.

  • Updated Jul 27, 2024
  • Jupyter Notebook

Split a dataset into subsets of specified sizes such that each subset preserves the original label distribution enhanced by stratification on UMAP-based pseudo-labels. This method ensures splits are balanced both by true labels and the data’s underlying manifold structure.

  • Updated Jul 19, 2025
  • Python

Predicting company bankruptcy using various machine learning models. The dataset is sourced from Kaggle: Company Bankruptcy Prediction.

  • Updated Aug 7, 2024
  • Python

In this project, I have used logistic regression, a supervised machine learning algorithm, to predict whether a person has diabetes or not based on various features such as age, blood pressure, glucose level, body mass index, etc. I have used Python and popular libraries such as Pandas, Scikit-Learn, and Matplotlib to perfom model building

  • Updated Jan 26, 2024
  • Jupyter Notebook

Julia package for "FDR Control via Data Splitting for Testing-after-Clustering (arXiv: 2410.06451)"

  • Updated Nov 6, 2024
  • Julia

Embedding-based method for splitting datasets semantically

  • Updated Jun 17, 2026
  • Jupyter Notebook

Comparative Analysis of Data Protection Mechanisms in Public Clouds

  • Updated Aug 29, 2023

A basic Python script to split a .dat file into individual sample files.

  • Updated Apr 21, 2021
  • Python

A simple PyTorch-based neural network that classifies student exam outcomes (Pass/Fail) using study hours and previous exam scores. Implements dataset splitting (train/val/test), mini-batch training, and evaluation with configurable hyperparameters.

  • Updated Oct 5, 2025
  • Python

Improve this page

Add a description, image, and links to the data-splitting topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the data-splitting topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL