Flexible HDF5 saving/loading and other data science tools from the University of Chicago

Last update: Dec 10, 2022

Overview

https://travis-ci.org/uchicago-cs/deepdish.svg?branch=master

https://img.shields.io/badge/license-BSD%203--Clause-blue.svg?style=flat

deepdish

Flexible HDF5 saving/loading and other data science tools from the University of Chicago. This repository also host a Deep Learning blog:

http://deepdish.io

Installation

pip install deepdish

Alternatively (if you have conda with the conda-forge channel):

conda install -c conda-forge deepdish

Main feature

The primary feature of deepdish is its ability to save and load all kinds of data as HDF5. It can save any Python data structure, offering the same ease of use as pickling or numpy.save. However, it improves by also offering:

Interoperability between languages (HDF5 is a popular standard)
Easy to inspect the content from the command line (using h5ls or our specialized tool ddls)
Highly compressed storage (thanks to a PyTables backend)
Native support for scipy sparse matrices and pandas DataFrame, Series and Panel
Ability to partially read files, even slices of arrays

An example:

import deepdish as dd

d = {
    'foo': np.ones((10, 20)),
    'sub': {
        'bar': 'a string',
        'baz': 1.23,
    },
}
dd.io.save('test.h5', d)

This can be reconstructed using dd.io.load('test.h5'), or inspected through the command line using either a standard tool:

$ h5ls test.h5
foo                      Dataset {10, 20}
sub                      Group

Or, better yet, our custom tool ddls (or python -m deepdish.io.ls):

$ ddls test.h5
/foo                       array (10, 20) [float64]
/sub                       dict
/sub/bar                   'a string' (8) [unicode]
/sub/baz                   1.23 [float64]

Documentation

http://deepdish.readthedocs.io/

Flexible HDF5 saving/loading and other data science tools from the University of Chicago

Related tags

Overview

deepdish

Installation

Main feature

Documentation

Owner

UChicago - Department of Computer Science

Using approximate bayesian posteriors in deep nets for active learning

PyPDC is a Python package for calculating asymptotic Partial Directed Coherence estimations for brain connectivity analysis.

Using Python to derive insights on particular Pokemon, Types, Generations, and Stats

A script to "SHUA" H1-2 map of Mercenaries mode of Hearthstone

An ETL framework + Monitoring UI/API (experimental project for learning purposes)

API>local_db>AWS_RDS - Disclaimer! All data used is for educational purposes only.

songplays datamart provide details about the musical taste of our customers and can help us to improve our recomendation system

A fast, flexible, and performant feature selection package for python.

High Dimensional Portfolio Selection with Cardinality Constraints

CleanX is an open source python library for exploring, cleaning and augmenting large datasets of X-rays, or certain other types of radiological images.

Full ELT process on GCP environment.

Project: Netflix Data Analysis and Visualization with Python

NumPy and Pandas interface to Big Data

Import, connect and transform data into Excel

Candlestick Pattern Recognition with Python and TA-Lib

Python script for transferring data between three drives in two separate stages

Data and code accompanying the paper Politics and Virality in the Time of Twitter

A data analysis using python and pandas to showcase trends in school performance.

Probabilistic reasoning and statistical analysis in TensorFlow

This module is used to create Convolutional AutoEncoders for Variational Data Assimilation