Subscribe to DSC Newsletter

Vincent Granville's Blog – February 2020 Archive (4)

State-of-the-Art Statistical Science to Tackle Famous Number Theory Conjectures

The methodology described here has broad applications, leading to new statistical tests, new type of ANOVA (analysis of variance), improved design of experiments, interesting fractional factorial designs, a better understanding of irrational numbers leading to cryptography, gaming and Fintech applications, and high quality random numbers generators (and when you really need them). It also features exact arithmetic / high performance computing and distributed algorithms to compute millions of…

Continue

Added by Vincent Granville on February 29, 2020 at 11:00pm — No Comments

Advanced Analytic Platforms – Changes in the Leaderboard 2020

Summary: The Gartner Magic Quadrant for Data Science and Machine Learning Platforms is just out the big news is how much more capable all the platforms have become.  Of course there are also some interesting winner and loser stories.

The Gartner Magic Quadrant for Data Science and Machine Learning Platforms is just out for 2020.  The really big news is how many excellent choices are now available.  In a remarkable move, the whole field…

Continue

Added by Vincent Granville on February 21, 2020 at 9:25am — No Comments

Sentiment Analysis with Naive Bayes and LSTM

In this notebook, we try to predict the positive (label 1) or negative (label 0) sentiment of the sentence. We use the UCI Sentiment Labelled Sentences Data Set.

Sentiment analysis is very useful in many areas. For example, it can be used for internet conversations moderation. Also, it is possible to predict ratings that users can assign to a certain product (food, household appliances, hotels,…

Continue

Added by Vincent Granville on February 19, 2020 at 8:42pm — No Comments

Common Errors in Machine Learning due to Poor Statistics Knowledge

Probably the worst error is thinking there is a correlation when that correlation is purely artificial. Take a data set with 100,000 variables, say with 10 observations. Compute all the (99,999 * 100,000) / 2 cross-correlations. You are almost guaranteed to find one above 0.999. This is best illustrated in may article How to Lie with P-values (also discussing…

Continue

Added by Vincent Granville on February 7, 2020 at 9:48am — No Comments

Monthly Archives

2020

2019

2018

2017

2016

2015

2014

2013

2012

2011

2010

2009

2008

On Data Science Central

© 2020   TechTarget, Inc.   Powered by

Badges  |  Report an Issue  |  Privacy Policy  |  Terms of Service