<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Machine-Learning on Gabriel Berardi</title><link>http://www.gabriel-berardi.com/tags/machine-learning/</link><description>Recent content in Machine-Learning on Gabriel Berardi</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Tue, 01 Dec 2020 00:00:00 +0000</lastBuildDate><atom:link href="http://www.gabriel-berardi.com/tags/machine-learning/index.xml" rel="self" type="application/rss+xml"/><item><title>Data Leakage in Machine Learning</title><link>http://www.gabriel-berardi.com/blog/data/2020-12-01-data-leakage-in-machine-learning/</link><pubDate>Tue, 01 Dec 2020 00:00:00 +0000</pubDate><guid>http://www.gabriel-berardi.com/blog/data/2020-12-01-data-leakage-in-machine-learning/</guid><description>&lt;p&gt;Recently, I read a thread on Twitter about several Machine Learning papers that contained severe cases of data leakage. The authors of the papers seemed unaware of this phenomenon and therefore trained models that performed exceptionally well. Unfortunately, this was mainly due to data leakage.&lt;/p&gt;
&lt;p&gt;Not many beginners are aware of this problem and in my opinion, not many courses emphasize this issue early enough. Therefore, I would like to tell you all the things you need to know about data leakage and some ways to prevent it in this post.&lt;/p&gt;</description></item><item><title>k-Nearest Neighbors</title><link>http://www.gabriel-berardi.com/blog/data/2020-10-01-k-nearest-neighbors/</link><pubDate>Thu, 01 Oct 2020 00:00:00 +0000</pubDate><guid>http://www.gabriel-berardi.com/blog/data/2020-10-01-k-nearest-neighbors/</guid><description>&lt;p&gt;k-Nearest Neighbors, or k-NN as I am going to call it from now on, is one of the easiest algorithms to solve classification tasks. It can be used for regression
problems as well, but I am going to focus on the more common use case of classification in this post.&lt;/p&gt;
&lt;p&gt;In a nutshell, k-NN will assign a new data point to the class that the majority of its k neighbors in the training set belongs to. Let&amp;rsquo;s use another coffee-related example to see how that works.&lt;/p&gt;</description></item><item><title>Detect Forged Banknotes with a Logistic Regression</title><link>http://www.gabriel-berardi.com/blog/data/2020-09-01-detect-forged-banknotes-with-a-logistic-regression/</link><pubDate>Tue, 01 Sep 2020 00:00:00 +0000</pubDate><guid>http://www.gabriel-berardi.com/blog/data/2020-09-01-detect-forged-banknotes-with-a-logistic-regression/</guid><description>&lt;p&gt;Counterfeit money is a serious problem for both individuals and businesses. Counterfeiters constantly find new ways and techniques to produce fake banknotes, that are essentially
indistinguishable from real money. At least for the human eye!&lt;/p&gt;
&lt;p&gt;Identifying forged banknotes is a typical example of a binary classification task in Machine Learning. If we have enough data of both real and forged banknotes, we can use this data to
train a model that can classify new banknotes as either real or fake.&lt;/p&gt;</description></item><item><title>Simple Facial Recognition with OpenCV</title><link>http://www.gabriel-berardi.com/blog/data/2020-08-01-simple-facial-recognition-with-opencv/</link><pubDate>Sat, 01 Aug 2020 00:00:00 +0000</pubDate><guid>http://www.gabriel-berardi.com/blog/data/2020-08-01-simple-facial-recognition-with-opencv/</guid><description>&lt;p&gt;Have you ever seen some cool applications of computer vision tools, like this the one below?&lt;/p&gt;
&lt;p&gt;Perhaps your phone&amp;rsquo;s camera can autofocus on faces, or maybe you have uploaded a photo on a social media platform and it automatically recognized the person on the image?&lt;/p&gt;
&lt;p&gt;These are facial recognition applications and they all rely on Machine Learning. In this post, we are going to use a very easy package called OpenCV to build our
own facial recognition program!&lt;/p&gt;</description></item><item><title>Linear and Logistic Regression</title><link>http://www.gabriel-berardi.com/blog/data/2020-07-01-linear-and-logistic-regression/</link><pubDate>Wed, 01 Jul 2020 00:00:00 +0000</pubDate><guid>http://www.gabriel-berardi.com/blog/data/2020-07-01-linear-and-logistic-regression/</guid><description>&lt;p&gt;Linear and Logistic regression are among the most elementary algorithms for supervised learning. Supervised Learning describes the situation where we deal with labeled data, which means that we have labeled inputs and a target variable.&lt;/p&gt;
&lt;p&gt;Despite the fact that both have the word &amp;ldquo;regression&amp;rdquo; in their name, only one of them is typically being used for solving regression problems!&lt;/p&gt;
&lt;p&gt;Let&amp;rsquo;s see how they work!&lt;/p&gt;
&lt;h2 id="linear-regression"&gt;Linear Regression&lt;/h2&gt;
&lt;p&gt;Linear regression is possibly the easiest, most intuitive way of making a quantitative prediction. The relationship between an independent and a dependent variable is assumed to be linear, meaning that the dependent variable can be predicted using a linear function of the independent variable. For example:&lt;/p&gt;</description></item><item><title>k-Means Clustering</title><link>http://www.gabriel-berardi.com/blog/data/2020-01-01-k-means-clustering/</link><pubDate>Wed, 01 Jan 2020 00:00:00 +0000</pubDate><guid>http://www.gabriel-berardi.com/blog/data/2020-01-01-k-means-clustering/</guid><description>&lt;p&gt;The k-means algorithm is used to divide unlabeled data into categories or classes, in order to draw useful conclusions from the resulting clusters.&lt;/p&gt;
&lt;p&gt;Let&amp;rsquo;s take a look at an imaginary dataset of n = 18 observations of different coffee brands. Note that we would never actually use the k-means algorithm on such a small dataset.&lt;/p&gt;
&lt;p&gt;We plot the price of the coffee vs. the rating obtained by customers:&lt;/p&gt;
&lt;p&gt;








&lt;a href="images/coffee-price-vs-rating-scatterplot.png" data-fancybox="post-images" data-caption="scatter plot of coffee price vs customer rating"&gt;
 &lt;img src="images/coffee-price-vs-rating-scatterplot.png" alt="scatter plot of coffee price vs customer rating" /&gt;
&lt;/a&gt;
&lt;/p&gt;</description></item></channel></rss>