Data Science Specialization

https://www.coursera.org/specializations/jhu-data-science

Launch Your Career in Data Science. A nine-course introduction to data science, developed and taught by leading professors.

Ask the right questions, manipulate data sets, and create visualizations to communicate results. This Specialization covers the concepts and tools you'll need throughout the entire data science pipeline, from asking the right kinds of questions to making inferences and publishing results. In the final Capstone Project, you’ll apply the skills learned by building a data product using real-world data. At completion, students will have a portfolio demonstrating their mastery of the material.

43 weeks - 367 hours

Certificate: https://www.coursera.org/account/accomplishments/specialization/WP3UYBMB4VUK

Course 1: The Data Scientist’s Toolbox

https://www.coursera.org/learn/data-scientists-tools

Neste curso você receberá uma introdução às ferramentas e idéias mais importantes da "caixa de ferramenta do cientista". O curso dará uma visão geral de dados, perguntas e ferramentas utilizadas pelo cientista para a análise de dados. Este curso possui duas partes. A primeira é uma parte introdutória e conceitual às ideias por trás da transformação de dados em conhecimento. A segunda parte dará uma intrudução prática às ferramentas que serão utilizadas durante o programa (version control, markdown, git, GitHub, R e RStudio).

4 weeks - 1-4 hours

Certificate: https://www.coursera.org/account/accomplishments/verify/2JAW2EURGG

Course 2: R Programming

https://www.coursera.org/learn/r-programming

In this course you will learn how to program in R and how to use R for effective data analysis. You will learn how to install and configure software necessary for a statistical programming environment and describe generic programming language concepts as they are implemented in a high-level statistical language. The course covers practical issues in statistical computing which includes programming in R, reading data into R, accessing R packages, writing R functions, debugging, profiling R code, and organizing and commenting R code. Topics in statistical data analysis will provide working examples.

4 weeks - 7-9 hours

Certificate: https://www.coursera.org/account/accomplishments/verify/XHBTABWSJG

Course 3: Getting and Cleaning Data

https://www.coursera.org/learn/data-cleaning

Before you can work with data you have to get some. This course will cover the basic ways that data can be obtained. The course will cover obtaining data from the web, from APIs, from databases and from colleagues in various formats. It will also cover the basics of data cleaning and how to make data “tidy”. Tidy data dramatically speed downstream data analysis tasks. The course will also cover the components of a complete data set including raw data, processing instructions, codebooks, and processed data. The course will cover the basics needed for collecting, cleaning, and sharing data.

4 weeks - 4-9 hours

Certificate: https://www.coursera.org/account/accomplishments/verify/NRT8JF72PM

Course 4: Exploratory Data Analysis

https://www.coursera.org/learn/exploratory-data-analysis

This course covers the essential exploratory techniques for summarizing data. These techniques are typically applied before formal modeling commences and can help inform the development of more complex statistical models. Exploratory techniques are also important for eliminating or sharpening potential hypotheses about the world that can be addressed by the data. We will cover in detail the plotting systems in R as well as some of the basic principles of constructing data graphics. We will also cover some of the common multivariate statistical techniques used to visualize high-dimensional data.

4 weeks - 4-9 hours

Certificate: https://www.coursera.org/account/accomplishments/verify/KPL23DFZFK

Course 5: Reproducible Research

https://www.coursera.org/learn/reproducible-research

This course focuses on the concepts and tools behind reporting modern data analyses in a reproducible manner. Reproducible research is the idea that data analyses, and more generally, scientific claims, are published with their data and software code so that others may verify the findings and build upon them. The need for reproducibility is increasing dramatically as data analyses become more complex, involving larger datasets and more sophisticated computations. Reproducibility allows for people to focus on the actual content of a data analysis, rather than on superficial details reported in a written summary. In addition, reproducibility makes an analysis more useful to others because the data and code that actually conducted the analysis are available. This course will focus on literate statistical analysis tools which allow one to publish data analyses in a single document that allows others to easily execute the same analysis to obtain the same results. In this course you will learn to write a document using R markdown, integrate live R code into a literate statistical program, compile R markdown documents using knitr and related tools, and organize a data analysis so that it is reproducible and accessible to others.

4 weeks - 4-9 hours

Certificate: https://www.coursera.org/account/accomplishments/verify/KPL23DFZFK

Course 6: Statistical Inference

https://www.coursera.org/learn/statistical-inference/home/info

Statistical inference is the process of drawing conclusions about populations or scientific truths from data. There are many modes of performing inference including statistical modeling, data oriented strategies and explicit use of designs and randomization in analyses. Furthermore, there are broad theories (frequentists, Bayesian, likelihood, design based, …) and numerous complexities (missing data, observed and unobserved confounding, biases) for performing inference. A practitioner can often be left in a debilitating maze of techniques, philosophies and nuance. This course presents the fundamentals of inference in a practical approach for getting things done. After taking this course, students will understand the broad directions of statistical inference and use this information for making informed choices in analyzing data.

4 weeks - 4-9 hours

Certificate: https://www.coursera.org/account/accomplishments/verify/K2X2KNTRUBM4

Course 7: Regression Models

https://www.coursera.org/learn/regression-models/home/info

Linear models, as their name implies, relates an outcome to a set of predictors of interest using linear assumptions. Regression models, a subset of linear models, are the most important statistical analysis tool in a data scientist’s toolkit. This course covers regression analysis, least squares and inference using regression models. Special cases of the regression model, ANOVA and ANCOVA will be covered as well. Analysis of residuals and variability will be investigated. The course will cover modern thinking on model selection and novel uses of regression models including scatterplot smoothing.

4 weeks - 4-9 hours

Certificate: https://www.coursera.org/account/accomplishments/verify/RTKGSW5MZ667

Course 8: Practical Machine Learning

https://www.coursera.org/learn/practical-machine-learning/home/info

One of the most common tasks performed by data scientists and data analysts are prediction and machine learning. This course will cover the basic components of building and applying prediction functions with an emphasis on practical applications. The course will provide basic grounding in concepts such as training and tests sets, overfitting, and error rates. The course will also introduce a range of model based and algorithmic machine learning methods including regression, classification trees, Naive Bayes, and random forests. The course will cover the complete process of building prediction functions including data collection, feature creation, algorithms, and evaluation.

4 weeks - 4-9 hours

Certificate: https://www.coursera.org/account/accomplishments/verify/PJHB97CD8GR4

Course 9: Developing Data Products

https://www.coursera.org/learn/data-products/home/info

A data product is the production output from a statistical analysis. Data products automate complex analysis tasks or use technology to expand the utility of a data informed model, algorithm or inference. This course covers the basics of creating data products using Shiny, R packages, and interactive graphics. The course will focus on the statistical fundamentals of creating a data product that can be used to tell a story about data to a mass audience.

4 weeks - 4-9 hours

Certificate: https://www.coursera.org/account/accomplishments/verify/U2H4NR5U59QS

Course 10: Data Science Capstone

https://www.coursera.org/learn/data-science-project/home/info

The capstone project class will allow students to create a usable/public data product that can be used to show your skills to potential employers. Projects will be drawn from real-world problems and will be conducted with industry, government, and academic partners.

7 weeks - 4-9 hours

Certificate: https://www.coursera.org/account/accomplishments/verify/S99TXTP52WCH

cardosop / data-science-specialization Goto Github PK