FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

spark · GitHub Topics · GitHub

#

Apache Spark

Apache Spark is an open source distributed general-purpose cluster-computing framework. It provides an interface for programming entire clusters with implicit data parallelism and fault tolerance.

Here are 10,340 public repositories matching this topic...

Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. Join the course here 👇🏼

  • Updated Aug 25, 2026
  • Jupyter Notebook

Apache Spark - A unified analytics engine for large-scale data processing

  • Updated Aug 28, 2026
  • Scala

Data science Python notebooks: Deep learning (TensorFlow, Theano, Caffe, Keras), scikit-learn, Kaggle, big data (Spark, Hadoop MapReduce, HDFS), matplotlib, pandas, NumPy, SciPy, Python essentials, AWS, and various command lines.

  • Updated Mar 20, 2024
  • Python

Make Your Company Data Driven. Connect to any data source, easily visualize, dashboard and share your data.

  • Updated Aug 18, 2026
  • Python

最新Docker容器技术,从真实案例中学习最佳实践!| Learn and understand Docker&Container technologies, with real DevOps practice!

  • Updated Aug 26, 2026
  • Go

大数据入门指南 ⭐

  • Updated Jan 5, 2024
  • Java

List of Data Science Cheatsheets to rule the world

  • Updated Jul 18, 2024

GUI for ChatGPT API and many LLMs. Supports agents, file-based QA, GPT finetuning and query with web search. All with a neat UI.

  • Updated Apr 30, 2026
  • Python

flink learning blog. http://www.54tianzhisheng.cn/ 含 Flink 入门、概念、原理、实战、性能调优、源码解析等内容。涉及 Flink Connector、Metrics、Library、DataStream API、Table API & SQL 等内容的学习案例,还有 Flink 落地应用的大型项目案例(PVUV、日志存储、百亿数据实时去重、监控告警)分享。欢迎大家支持我的专栏《大数据实时计算引擎 Flink 实战与性能优化》

  • Updated May 6, 2026
  • Java

【大厂面试专栏】一份Java程序员需要的技术指南,这里有面试题、系统架构、职场锦囊、主流中间件等,让你成为更牛的自己!

  • Updated Jul 21, 2025

Distributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet.

  • Updated Jul 29, 2026
  • Python

Suite of tools for deploying and training deep learning models using the JVM. Highlights include model import for keras, tensorflow, and onnx/pytorch, a modular and tiny c++ library for running math code and a java based math library on top of the core c++ library. Also includes samediff: a pytorch/tensorflow like library for running deep learn...

  • Updated Aug 28, 2026
  • Java

专注大数据学习面试,大数据成神之路开启。Flink/Spark/Hadoop/Hbase/Hive...

  • Updated Aug 7, 2023

An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Trino, and Hive and APIs

  • Updated Aug 28, 2026
  • Scala

🧙 Build, run, and manage data pipelines for integrating and transforming data.

  • Updated Aug 13, 2026
  • Python

H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc.

  • Updated Aug 26, 2026
  • Jupyter Notebook

Alluxio, data orchestration for analytics and machine learning in the cloud

  • Updated Apr 29, 2025
  • Java

A Flexible and Powerful Parameter Server for large-scale machine learning

  • Updated Jul 26, 2026
  • Java

Web-based notebook that enables data-driven, interactive data analytics and collaborative documents with SQL, Scala and more.

  • Updated Aug 26, 2026
  • Java

Created by Matei Zaharia

Released May 26, 2014

Followers
440 followers
Repository
apache/spark
Website
github.com/topics/spark
Wikipedia
Wikipedia

Related topics

hadoop scala

Back | FazBrowse Home | New Git URL