Level I Β· Quantitative

Learning Module 11
Big Data Techniques

Key Outcomes Summary & Practice Problems

Learning Outcomes

What you must be able to do

Curriculum Year: 2026

LOS 1

Describe aspects of "fintech" that are directly relevant for the gathering and analyzing of financial data.

LOS 2

Describe Big Data, artificial intelligence, and machine learning β€” including types of ML and data structures.

LOS 3

Describe applications of Big Data and data science to investment management, including NLP, risk analysis, and algorithmic trading.

1 Β· Fintech & Big Data

Fintech β€” the meeting of finance and technology β€” is transforming investment management through Big Data, AI, and machine learning. Two areas most relevant to quantitative analysis: (1) analysis of large datasets including alternative data from non-traditional sources, and (2) analytical tools such as AI that identify complex, non-linear relationships beyond traditional statistics.

Big Data refers to the vast amount of information generated by industry, governments, individuals, and electronic devices since the late 1990s. It encompasses both traditional data (stock prices, company financials, government statistics) and alternative data β€” data from social media, sensors, company exhaust, and the Internet of Things that arise from the normal course of modern digital life.

THE FOUR VS OF BIG DATA

1ST V
Volume

Massive amounts of data β€” millions to billions of data points. From megabytes β†’ gigabytes β†’ terabytes β†’ petabytes.

2ND V
Velocity

Speed and frequency of data recording and transmission. Real-time or near-real-time data is now the norm in many areas.

3RD V
Variety

Data from many sources in diverse formats: structured (SQL tables), semistructured (HTML), and unstructured (video, social media).

4TH V
Veracity

Credibility and reliability of data sources. Critical for inference/prediction. Big Data amplifies the challenge of quality vs quantity.

Data Structure Types

Type

Description

Examples

Storage

Structured

Organised in tables; each field same type of information

SQL tables, spreadsheets, financial databases

SQL databases

Semi-structured

Attributes of both structured and unstructured; has some organisation

HTML code, XML files, JSON data

NoSQL or special parsers

Unstructured

Disparate, cannot be represented in tabular form; often requires custom processing

Social media posts, emails, voice recordings, videos, satellite imagery

NoSQL databases

    • Alternative data sources: data from electronic devices, social media, sensor networks, and company exhaust. Increasingly used to generate alpha, reduce losses, and gain real-time insights not available from traditional quarterly/annual filings.

    • Big Data challenges: selection bias, missing data, outliers, and questions of whether the volume and type of data are appropriate for the analysis. Data must be sourced, cleansed, and organised before use β€” a significant hurdle with unstructured alternative data.

    • Legal and ethical issues: scraping web data may capture personally protected information. Best practices are still developing, and national regulators take different approaches. Investment professionals must be aware of compliance requirements.

2 Β· Three Sources of Alternative Data

SOURCE 1

πŸ‘€ Individuals

Text, video, photo, audio β€” and digital actions like website clicks or time on page. Typically unstructured. Volume growing dramatically with online participation.

Social media Β· News reviews Β· Web searches Β· Personal digital trails

SOURCE 2

🏒 Business Processes

Information flows from corporations and public entities. Tends to be structured data. Can be leading or real-time indicators vs traditional lagging metrics.

Credit card data Β· Corporate exhaust Β· Supply chain Β· Point-of-sale scanner data Β· Banking records

SOURCE 3

πŸ“‘ Sensors

Smart phones, cameras, RFID chips, satellites connected via wireless networks. Often unstructured. Orders of magnitude larger than other data streams. Growing exponentially via IoT.

Satellite imagery Β· Geolocation Β· Internet of Things (IoT) Β· Traffic patterns Β· Shipping cargo data

Investment applications of alternative data: Satellite imagery of retail parking lots β†’ foot traffic data ahead of earnings; shipping activity β†’ supply chain indicators; agricultural satellite data β†’ crop yield forecasts; social media sentiment β†’ predictive signals for stock returns and IPO performance. Alternative data can identify factors affecting security prices, improve asset selection, optimize trade execution, and uncover trends before they appear in traditional financial reports.

3 Β· Artificial Intelligence & Machine Learning

Artificial Intelligence (AI) computer systems perform tasks that traditionally required human intelligence at levels comparable or superior to humans. Machine Learning (ML) is a subset of AI that extracts knowledge from large datasets without assuming any underlying probability distribution. ML goal: "find the pattern, apply the pattern."

ML training process β€” three dataset splits:
1. Training dataset β€” algorithm learns input-output relationships from historical patterns
2. Validation dataset β€” relationships are validated and the model tuned
3. Test dataset β€” model's ability to predict on new, unseen data is evaluated

Once trained, validated, and tested, the ML model predicts outcomes on other datasets. ML still requires human judgment in understanding data and selecting appropriate techniques.

THREE CLASSES OF MACHINE LEARNING

πŸ“Œ SUPERVISED LEARNING

Computers learn to model relationships based on labeled training data β€” inputs AND outputs are both identified for the algorithm. After learning, trained algorithms model or predict outcomes for new datasets.

Finance examples: identifying best signal to forecast stock returns; predicting whether a local market will be up/down/flat.

πŸ” UNSUPERVISED LEARNING

Computers receive only data β€” no labeled outputs. The algorithm seeks to describe the data and their structure without any predefined categories.

Finance examples: grouping companies into peer groups by characteristics (rather than standard sectors); clustering bonds by risk profile.

🧠 DEEP LEARNING

Uses neural networks with many hidden layers for multistage, non-linear data processing. Can use supervised or unsupervised approaches. Builds understanding from simple to complex concepts in layers.

Applications: image recognition, speech recognition, pattern recognition in financial data.

    • Overfitting: the model learns training data too precisely, treating noise as true parameters. The overfitted model cannot accurately predict on new (out-of-sample) datasets β€” may be too complex.

    • Underfitting: the model treats true parameters as noise and fails to recognise real patterns in training data. Results in a model that is too simplistic and misses underlying patterns.

    • "Black box" problem: because ML algorithms are not explicitly programmed, their outcomes may not be fully understood or explainable β€” a significant concern for regulatory compliance and investment governance.

    • Neural networks have existed since 1958; used for forecasting and pattern recognition. Improvements in underlying algorithms now enable better image, pattern, and speech recognition with less computing power.

    • Expert systems (early AI): "if–then" rules designed to simulate expert human judgment. Later replaced by more sophisticated ML techniques that learn from data rather than following hard-coded rules.

4 Β· Data Science & Investment Management Applications

Data science is an interdisciplinary field combining computer science (including ML), statistics, and other disciplines to extract information from Big Data. Data scientists handle the full pipeline from raw data to actionable insight.

FIVE DATA PROCESSING METHODS

1
πŸ“₯
CAPTURE

How data are collected and transformed into usable format. Low-latency systems for real-time trading; high-latency for slower analysis.

2
🧹
CURATION

Ensuring data quality through data cleaning β€” detecting errors, inaccuracies, and making adjustments for missing data.

3
πŸ—„οΈ
STORAGE

How data are recorded, archived, and accessed. Key considerations: structured vs unstructured, and latency requirements.

4
πŸ”Ž
SEARCH

How to query data. Big Data requires advanced applications capable of examining large quantities to locate requested content.

5
πŸ”„
TRANSFER

How data move from source or storage to analytical tools β€” e.g., a stock exchange's real-time price feed.

6
πŸ”Ž
SEARCH

How to query data. Big Data requires advanced applications capable of examining large quantities to locate requested content.

7
πŸ”„
TRANSFER

How data move from source or storage to analytical tools β€” e.g., a stock exchange's real-time price feed.

KEY INVESTMENT MANAGEMENT APPLICATIONS

Application

Description

Finance Use Cases

Text Analytics

Computer programs analyse and derive meaning from large, unstructured text/voice datasets. Includes lexical analysis (word frequency) and pattern recognition.

Company filings Β· Earnings call transcripts Β· Social media Β· Consumer sentiment indicators

Natural Language Processing (NLP)

AI-powered field at the intersection of CS, AI, and linguistics. Automates translation, speech recognition, text mining, sentiment analysis, and topic analysis.

Analyst sentiment tagging ahead of recommendation changes Β· Central bank communications analysis Β· EPS forecast nuance detection Β· Compliance monitoring

Risk Analysis

ML techniques applied to large datasets to identify and quantify risks beyond traditional models. Can detect complex, non-linear risk relationships.

Portfolio risk modelling Β· Credit risk assessment Β· Fraud detection (neural networks in credit card systems since 1980s)

Algorithmic Trading

Computer programs execute trading decisions at speeds impossible for humans, using pre-defined rules or ML-based logic. Requires low-latency systems.

High-frequency trading (C/C++) Β· Market trend prediction Β· Merger outcome prediction Β· Execution optimisation

Data Visualisation

Formatting, displaying, and summarising data in graphical form. Traditional: tables and charts. Non-traditional: 3D interactive graphics, heat maps, tree diagrams, tag clouds, mind maps.

Tag cloud β€” word frequency from analyst reports; Mind map β€” relationship between concepts; Network graphs for correlation structures

ILLUSTRATIVE TAG CLOUD β€” WORD FREQUENCY IN BIG DATA TECHNIQUES
data ML learning AI algorithms techniques model training patterns neural relationships fintech predict inputs outputs deep identify
Words sized by frequency Β· Tag clouds make text data visually accessible Β· A key data visualization tool in Big Data analytics
    • NLP & analyst sentiment: NLP assigns sentiment ratings (very negative β†’ very positive) to analyst commentary β€” detecting sentiment shifts ahead of formal recommendation changes. Scales across thousands of companies globally.

    • Central bank communications: NLP analyses Fed/ECB transcripts and communications for subtle messaging about interest rate policy, inflation expectations, and aggregate output β€” valuable for fixed income and macro strategies.

    • Programming languages: Python (most common, open-source, fintech-friendly); R (statistical analysis, ML packages); Java (internet applications); C/C++ (speed-critical, used in HFT); Excel VBA (automation, bridging programming and manual processing).

    • Databases: SQL (structured data in tables, server-based); SQLite (structured, embedded in programs, common for mobile apps); NoSQL (unstructured data that cannot be organized in rows/columns).