A Data Science Central Community
The new version of a book by Jeffrey Stanton from Syracuse Iniversity School of Information Studies, , is now available for free download. The book, developed for Syracuse's Certificate for Data Science, is available under a Creative Commons License.
The book begins with the following definition of Data Science:
Data Science refers to an emerging area of work concerned with the collection, preparation, analysis, visualization, management and preservation of large collections of information. Although the name Data Science seems to connect most strongly with areas such as databases and computer science, many different kinds of skills - including non-mathematical skills, are needed.
Throughout the book, you'll find many examples of data science applications implemented in the R language. For R beginners a Getting Started with R chapter is included, but it does get into some fairly in-depth topics including sentiment analysis of Twitter data, working with data in Hadoop via RHadoop, and creating information maps. R code is sprinkled liberally for your own use, and available to download (also under an open-source license) from GitHub.