News and Blog

Machine Learning Algorithms: Semi-supervised Learning and Reinforcement Learning

Giribone
AI in banking managementNews

Machine Learning Algorithms: Semi-supervised Learning and Reinforcement Learning

Edited by Pier Giuseppe Giribone

The previous two articles in this column discussed the computational paradigms associated with supervised and unsupervised learning of a machine learning system.

This article continues the excursus, illustrating the last two classification criteria, based on the types of learning, in particular we will focus on semi-supervised or “weak” learning (semi-supervised learning) and “reinforcement learning” (Reinforcement learning).

Since organizing a large amount of data is challenging and expensive, both in terms of time and money, an intelligent systems designer frequently finds himself having to manage a large number of unclassified (or rather unlabeled) instances and a limited number of labeled ones.

Up until now, these have been machine learning algorithms, which can process either labeled data using supervised learning algorithms or unlabeled data using unsupervised learning algorithms. Therefore, a mixed data set, such as the one described above, could not be processed in its entirety.

Machine learning algorithms that can handle data that is only partially labeled are called semi-supervised learning.

These should be understood as a combination of unsupervised and supervised algorithms: for this characteristic they are identified in the literature as “hybrids”.

To understand how a semi-supervised algorithm works, an intuitive example is provided below, aimed at understanding its applicability in a banking context.

Cloud photo hosting services, such as Google Photos, can be good examples. When photos of a family outing are uploaded to this service, the algorithm is able to automatically identify that the same person A appears in photos 1, 4, and 7, while another person B appears in photos 2, 4, and 8. This first step is clearly performed by an algorithm characterized by unsupervised learning (clustering). Now all the system needs to know, to perform efficient organization, is who person A and person B are. The user then labels person A once with their name, and the label will automatically be transferred to all photos containing person A. The same procedure is repeated for all the people recognized during the clustering phase. This labeling makes it easier, for example, to search for a person by name within the photos uploaded to the service.

If a hybrid approach had not been used, all the people in each individual photo uploaded to the online archive system would have had to be manually labeled.

The use of a semi-supervised algorithm therefore made it possible to achieve the goal of data labeling effectively and efficiently.

In the banking sector, thanks to the increasingly advanced digitalization of financial transactions, repositories containing enormous amounts of data have been created, making organizing individual records using specific labels or indexed keys a task that is in fact too demanding, time-consuming, and costly.

As a result, the machine learning algorithm designer finds himself working with a small amount of data organized by labels mixed with a large amount of disorganized data (without labels).

The application of a semi-supervised algorithm, due to its hybrid nature, could represent a good compromise to be able to perform automatic labeling on the available data on the one hand and to be able to apply statistical inferences in the presence of mixed data (i.e. with and without labels) on the other.

Reinforcement Learning (RL) is a conceptually different learning approach than the previous ones. The learning system, called an agent in this context, is able to observe and interact with the surrounding environment, choose to perform actions, and, upon completion, receive either positive (positive reward) or negative (negative reward or penalty) feedback.

The policy defines the best action the agent should take in a given context.

The feedback received allows the agent to learn autonomously, starting from the choices made, selecting the best strategy (called policy) that allows him to maximize positive rewards over time.

The operating principle of RL can be extended to various applications:

– Robotics: The agent can be the program that controls a robot. In this case, the surrounding environment is the real world; the agent observes the environment using a set of sensors, and its actions consist of sending signals to activate the motors. It can be programmed to receive positive feedback if the robot reaches its target destination, and negative feedback if it wastes time or goes in the wrong direction.

– Video game: The agent could be the program that controls Ms. Pac-Man in the arcade game of the same name. In this case, the environment is a simulation of the Atari game, the actions are the possible positions the joystick can assume, the observations are the screenshots, and the rewards are the game points.

– Boardgame: similarly to the previous case, the agent can be the program that plays the ancient abstract board game "Go." DeepMind's AlphaGo program is a very popular example, as it beat world champion Ke Jie in May 2017. It learned the winning strategy (policy) by analyzing millions of games and then playing against itself numerous times. Obviously, the final test was conducted by disabling the learning phase during the match: the agent played fairly, applying only what it had learned from its previous training.

– Home automation: the agent doesn't necessarily have to have physical or virtual control over an object's movement. For example, it could be a smart thermostat, which receives positive rewards whenever it approaches the target temperature and saves energy and generates negative rewards when the human needs to adjust the temperature, so that the agent can anticipate this need.

– Algotrading: the agent can observe stock market prices and decide how much to buy or sell at any given time. The rewards in this case are obviously the realized gains or losses.

This article concludes the description of the classification of Machine Learning algorithms based on the amount and type of supervision they are subjected to during training (training).

The next article will discuss other allocation criteria, focusing on different learning modalities: “batch” versus “online” learning and “instance-based” versus “model-based” learning.

Select the fields to be shown. Others want to be hidden. Drag and drop to rearrange the order.
  • Image
  • SKU
  • Rating
  • Price
  • Stock
  • Availability
  • Add to Cart
  • Description
  • Content
  • Weight
  • Size
  • Product information
Click outside to hide the comparison bar
Compare