> For the complete documentation index, see [llms.txt](https://dshub.gitbook.io/ds-hub/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://dshub.gitbook.io/ds-hub/machine-learning/fundamentals.md).

# Fundamentals

Machine learning focuses on the development of algorithms and models that enable computers to learn and make predictions or decisions without being explicitly programmed. In essence, machine learning systems use data to improve their performance on a specific task or problem over time. The fundamental idea behind machine learning is to give computers the ability to learn from experience, much like humans do, but at a much faster and data-driven scale.

Machine learning can be categorized into different types based on the learning process:

* [**Supervised Learning**](/ds-hub/machine-learning/fundamentals/supervised-learning.md)**:** In supervised learning, the model learns from labeled data and makes predictions or classifications based on the labels.
* [**Unsupervised Learning**](/ds-hub/machine-learning/fundamentals/unsupervised-learning.md)**:** Unsupervised learning involves learning patterns and structures in data without explicit labels. Common techniques include clustering and dimensionality reduction.
* **Reinforcement Learning:** Reinforcement learning is used in scenarios where an agent interacts with an environment and learns by receiving rewards or penalties for its actions.

## Supervised vs Unsupervised Learning

Below is a comparison of some of the major differences between the two machine learning techniques:

<table><thead><tr><th width="172">Aspect</th><th width="372">Supervised Learning</th><th width="659">Unsupervised Learning</th></tr></thead><tbody><tr><td><strong>Definition</strong></td><td>Involves training a model on labeled data, where both the input features and the corresponding output labels are provided.</td><td>Involves training a model on unlabeled data, where only input features are provided, and the model must find patterns or structures within the data.</td></tr><tr><td><strong>Goal</strong></td><td>Predict or classify new data points based on past examples.</td><td>Discover underlying structures such as clusters or patterns.</td></tr><tr><td><strong>Output</strong></td><td>Predictive model or classification label for each input sample.</td><td>Clusters, patterns, or dimensionality reduction without labels.</td></tr><tr><td><strong>Examples</strong></td><td>Linear regression, decision trees, random forests, SVM.</td><td>K-means clustering, DBSCAN, PCA.</td></tr><tr><td><strong>Data Dependency</strong></td><td>Requires labeled data, which can be costly and time-consuming to collect.</td><td>Works with raw, unannotated data, making it more flexible.</td></tr><tr><td><strong>Accuracy</strong></td><td>Generally higher accuracy with labeled data and well-tuned models.</td><td>Accuracy can be harder to quantify, as there are no labels for validation.</td></tr><tr><td><strong>Complexity</strong></td><td>Can be computationally intensive, especially with large datasets or complex models.</td><td>Often simpler in terms of model complexity, but can still be computationally expensive.</td></tr></tbody></table>
