Distributed deep learning on Hadoop and Spark clusters.

Last update: Dec 28, 2022

Related tags

Overview

Note: we're lovingly marking this project as Archived since we're no longer supporting it. You are welcome to read the code and fork your own version of it and continue to use this code under the terms of the project license.

CaffeOnSpark

What's CaffeOnSpark?

CaffeOnSpark brings deep learning to Hadoop and Spark clusters. By combining salient features from deep learning framework Caffe and big-data frameworks Apache Spark and Apache Hadoop, CaffeOnSpark enables distributed deep learning on a cluster of GPU and CPU servers.

As a distributed extension of Caffe, CaffeOnSpark supports neural network model training, testing, and feature extraction. Caffe users can now perform distributed learning using their existing LMDB data files and minorly adjusted network configuration (as illustrated).

CaffeOnSpark is a Spark package for deep learning. It is complementary to non-deep learning libraries MLlib and Spark SQL. CaffeOnSpark's Scala API provides Spark applications with an easy mechanism to invoke deep learning (see sample) over distributed datasets.

CaffeOnSpark was developed by Yahoo for large-scale distributed deep learning on our Hadoop clusters in Yahoo's private cloud. It's been in use by Yahoo for image search, content classification and several other use cases.

Why CaffeOnSpark?

CaffeOnSpark provides some important benefits (see our blog) over alternative deep learning solutions.

It enables model training, test and feature extraction directly on Hadoop datasets stored in HDFS on Hadoop clusters.
It turns your Hadoop or Spark cluster(s) into a powerful platform for deep learning, without the need to set up a new dedicated cluster for deep learning separately.
Server-to-server direct communication (Ethernet or InfiniBand) achieves faster learning and eliminates scalability bottleneck.
Caffe users' existing datasets (e.g. LMDB) and configurations could be applied for distributed learning without any conversion needed.
High-level API empowers Spark applications to easily conduct deep learning.
Incremental learning is supported to leverage previously trained models or snapshots.
Additional data formats and network interfaces could be easily added.
It can be easily deployed on public cloud (ex. AWS EC2) or a private cloud.

Using CaffeOnSpark

Please check CaffeOnSpark wiki site for detailed documentations such as building instruction, API reference and getting started guides for standalone cluster and AWS EC2 cluster.

Batch sizes specified in prototxt files are per device.
Memory layers should not be shared among GPUs, and thus "share_in_parallel: false" is required for layer configuration.

Building for Spark 2.X

CaffeOnSpark supports both Spark 1.x and 2.x. For Spark 2.0, our default settings are:

spark-2.0.0
hadoop-2.7.1
scala-2.11.7 You may want to adjust them in caffe-grid/pom.xml.

Mailing List

Please join CaffeOnSpark user group for discussions and questions.

License

The use and distribution terms for this software are covered by the Apache 2.0 license. See LICENSE file for terms.

Distributed deep learning on Hadoop and Spark clusters.

Related tags

Overview

Note: we're lovingly marking this project as Archived since we're no longer supporting it. You are welcome to read the code and fork your own version of it and continue to use this code under the terms of the project license.

CaffeOnSpark

What's CaffeOnSpark?

Why CaffeOnSpark?

Using CaffeOnSpark

Building for Spark 2.X

Mailing List

License

Owner

Yahoo

Simplify stop motion animation with machine learning.

Repositório para o #alurachallengedatascience1

Distributed scikit-learn meta-estimators in PySpark

This is a curated list of medical data for machine learning

Extreme Learning Machine implementation in Python

Tools for mathematical optimization region

Kubeflow is a machine learning (ML) toolkit that is dedicated to making deployments of ML workflows on Kubernetes simple, portable, and scalable.

Estudos e projetos feitos com PySpark.

An AutoML survey focusing on practical systems.

A Python library for detecting patterns and anomalies in massive datasets using the Matrix Profile

Apache Liminal is an end-to-end platform for data engineers & scientists, allowing them to build, train and deploy machine learning models in a robust and agile way

A game theoretic approach to explain the output of any machine learning model.

Send rockets to Mars with artificial intelligence(Genetic algorithm) in python.

Quantum Machine Learning

The code from the Machine Learning Bookcamp book and a free course based on the book

Predict the output which should give a fair idea about the chances of admission for a student for a particular university

Pandas-method-chaining is a plugin for flake8 that provides method chaining linting for pandas code

Decentralized deep learning in PyTorch. Built to train models on thousands of volunteers across the world.

Getting Profit and Loss Make Easy From Binance

XGBoost-Ray is a distributed backend for XGBoost, built on top of distributed computing framework Ray.