created & maintained by @clarecorthell, founding partner of Luminant Data Science Consulting
Contents
- [The Open-Source Data Science Masters](#the-open-source-data-science-masters) - [Contents](#contents) - [The Internet is Your Oyster](#the-internet-is-your-oyster) - [The Motivation](#the-motivation) - [An Academic Shortfall](#an-academic-shortfall) - [Ready?](#ready) - [The Open Source Data Science Curriculum](#the-open-source-data-science-curriculum) - [A Note About Direction](#a-note-about-direction) - [Math](#math) - [Computing](#computing) - [Data Analysis](#data-analysis) - [Data Communication and Design](#data-communication-and-design) - [Python (Learning)](#python-learning) - [Python (Libraries)](#python-libraries) - [Datasets are now here](#datasets-are-now-here) - [R resources are now here](#r-resources-are-now-here) - [Data Science as a Profession](#data-science-as-a-profession) - [Capstone Project](#capstone-project) - [Resources](#resources) - [Read](#read) - [Watch & Listen](#watch--listen) - [Learn](#learn) - [Notation](#notation) - [Contribute](#contribute)1) Start Here
Intro to Data Science / UW Videos
- Topics: Python NLP on Twitter API, Distributed Computing Paradigm, MapReduce/Hadoop & Pig Script, SQL/NoSQL, Relational Algebra, Experiment design, Statistics, Graphs, Amazon EC2, Visualization.
Data Science / Harvard Videos & Course
- Topics: Data wrangling, data management, exploratory data analysis to generate hypotheses and intuition, prediction based on statistical methods such as regression and classification, communication of results through visualization, stories, and summaries.
Data Science with Open Source Tools Book $27
- Topics: Visualizing Data, Estimation, Models from Scaling Arguments, Arguments from Probability Models, What you Really Need to Know about Classical Statistics, Data Mining, Clustering, PCA, Map/Reduce, Predictive Analytics
- Example Code in: R, Python, Sage, C, Gnu Scientific Library
Human impact is a first-class concern when building machine intelligence technology. When we build products, we deduce patterns and then reinforce them in the world. Ethics in any Engineering concerns understanding the sociotechnological impact of the products and services we are bringing to bear in the human world -- and whether they are reinforcing a future we all want to live in.
2) Math and Problem-Solving
- Linear Algebra Khan Academy / Videos
- Linear Algebra / Levandosky Stanford / Book
$10
- Linear Programming (Math 407) University of Washington / Course
- The Manga Guide to Linear Algebra Book
$19
- An Intuitive Guide to Linear Algebra Better Explained / Article
- A Programmer's Intuition for Matrix Multiplication Better Explained / Article
- Vector Calculus: Understanding the Cross Product Better Explained / Article
- Vector Calculus: Understanding the Dot Product Better Explained / Article
- Convex Optimization / Boyd Stanford / Lectures / Book
- Stats in a Nutshell Book
$29
- Think Stats: Probability and Statistics for Programmers Digital & Book
$25
- Think Bayes Digital & Book
$25
- Differential Equations in Data Science Python Tutorial
- Problem-Solving Heuristics "How To Solve It" Polya / Book
$10
3) Databases, Distributed Computing, and Data Design
Get your environment up and running with the [Data Science Toolbox](http://bit.ly/datascitoolbox)- *See Intro to Data Science UW / Lectures on MapReduce
- Intro to Hadoop and MapReduce Cloudera / Udacity Course *includes select free excerpts of Hadoop: The Definitive Guide Book
$29
- Introduction to Databases Stanford / Online Course
- SQL School Mode Analytics / Tutorials
- SQL Tutorials SQLZOO / Tutorials
- Mining Massive Data Sets / Stanford Coursera & Digital & Book
$58
- Mining The Social Web Book
$30
- Introduction to Information Retrieval / Stanford Digital & Book
$56
How does the real world get translated into data? How should one structure that data to make it understandable and usable? Extends beyond database design to usability of schemas and models.
OSDSM Specialization: Web Scraping & Crawling
4) Applied Data Science: Beginner
Machine LearningFoundational & Theoretical
- Machine Learning Ng Stanford / Coursera & Stanford CS 229
- A Course in Machine Learning UMD / Digital Book
- The Elements of Statistical Learning / Stanford Digital & Book
$80
& Study Group - Machine Learning Caltech / Edx
Practical
- Programming Collective Intelligence Book
$27
- Machine Learning for Hackers ipynb / digital book
- Intro to scikit-learn, SciPy2013 youtube tutorials
- Probabilistic Programming and Bayesian Methods for Hackers Github / Tutorials
- Probabilistic Graphical Models Stanford / Coursera
- Neural Networks Andrej Karpathy / Python Walkthrough
- Neural Networks U Toronto / Coursera
- Deep Learning for Natural Language Processing CS224d Stanford
5) Applied Data Science: Intermediate
- Social and Economic Networks: Models and Analysis / Stanford / Coursera
- Social Network Analysis for Startups Book
$22
- From Languages to Information / Stanford CS147 Materials
- NLP with Python (NLTK library) Digital, Book
$36
- How to Write a Spelling Correcter / Norvig (Tutorial)[http://norvig.com/spell-correct.html]
One of the "unteachable" skills of data science is an intuition for analysis. What constitutes valuable, achievable, and well-designed analysis is extremely dependent on context and ends at hand.
- Big Data Analysis with Twitter UC Berkeley / Lectures
- Exploratory Data Analysis Tukey / Book
$81
7 and 8: Data Communication and Design
7) Data Visualization
_Data Visualization and Communication_ * The Truthful Art: Data, Charts, and Maps for Communication [Cairo / Book ```$21```](http://amzn.to/1UydGAc)Theoretical Design of Information
- Envisioning Information Tufte / Book
$36
- The Visual Display of Quantitative Information Tufte / Book
$27
Applied Design of Information
- Information Dashboard Design: Displaying Data for At-a-Glance Monitoring Stephen Few / Book
$29
Theoretical Courses / Design & Visualization
- Data Visualization University of Washington / Slides & Resources
- Berkeley's Viz Class UC Berkeley / Course Docs
- Rice University's Data Viz class Rice University / Slides
Practical Visualization Resources
- D3 Library / Scott Murray Blog / Tutorials
- Interactive Data Visualization for the Web / Scott Murray Online Book & Book
$26
OSDSM Specialization: Data Journalism
8) Python Libraries, Packages, and APIs
Installing Basic Packages [Python, virtualenv, NumPy, SciPy, matplotlib and IPython ](http://bit.ly/scientific-py-install) & [Using Python Scientifically](http://bit.ly/lecture-scipy)Command Line Install Script for Scientific Python Packages
- numpy Tutorial / Stanford CS231N
- Pandas Cookbook (data structure library)
More Libraries can be found in the "awesome machine learning" repo & in related specializations
- Flexible and powerful data analysis / manipulation library with labeled data structures objects, statistical functions, etc pandas & Tutorials Python for Data Analysis / Book
- scikit-learn - Tools for Data Mining & Analysis
- networkx - Network Modeling & Viz
- PyMC - Bayesian Inference & Markov Chain Monte Carlo sampling toolkit
- Statsmodels - Python module that allows users to explore data, estimate statistical models, and perform statistical tests
- PyMVPA - Multivariate Pattern Analysis in Python
- NLTK - Natural Language Toolkit
- Gensim - Python library for topic modeling, document indexing and similarity retrieval with large corpora. Target audience is the natural language processing (NLP) and information retrieval (IR) community.
- twython - Python wrapper for the Twitter API
- matplotlib - well-integrated with analysis and data manipulation packages like numpy and pandas
- Seaborn - a high-level statistical visualization package built on top of matplotlib
9) Capstone Project
* Capstone Analysis of Your Own Design; [Quora](http://bit.ly/quora-toyproblems)'s Idea Compendium * Healthcare Twitter Analysis [Coursolve & UW Data Science](http://bit.ly/project-healthcare-twitter-analysis) * Analyze your LinkedIn Network [Generate & Download Adjacency Matrix](http://socilab.com/)Resources
#### Read * [DataTau](http://bit.ly/datatau) - The "Hacker News" of Data Science * [The Signal and The Noise - Nate Silver ```$15```](http://amzn.to/1hoxQoG) - Bestseller Pop Sci * [Zipfian Academy's List of Resources](http://bit.ly/1qoF1We) * [A Software Engineer's Guide to Getting Started with Data Science](http://bit.ly/1jwgV4p) * [Data Scientist Interviews / Metamarkets](http://bit.ly/1r1tJot) * [/r/MachineLearning](http://bit.ly/1uANaEM)- The Life of a Data Scientist / Josh Wills
- The Talking Machines - Podcast about Machine Learning
- What Data Science Is / Hilary Mason
- Data Science in IPython Notebooks (Linear Regression, Logistic Regression, Random Forests, K-Means Clustering)
- A Gallery of Interesting IPython Notebooks - Pandas for Data Analysis
Datasets are now here
R resources are now here
- Doing Data Science: Straight Talk from the Frontline O'Reilly / Book
$25
- The Data Science Handbook: Advice and Insights from 25 Amazing Data Scientists Book
$22