Publication in the Diário da República: Despacho n.º 13495/2022 - 18/11/2022
10 ECTS; 1º Ano, 1º Semestre, 30,0 PL + 30,0 TP + 30,0 OT , Cód. 390913.
Lecturer
- Renato Eduardo Silva Panda (1)(2)
(1) Lead Professor
(2) Teaching Professor
Prerequisites
NA
Objectives
This course develops skills in data acquisition, preparation, storage, querying and analysis, connecting data science with data engineering and large-scale processing. By the end of the course, students should be able to:
1. Explain the data project lifecycle, the characteristics of Big Data and the ethical implications of collecting and using data.
2. Configure reproducible development environments and use Python and its libraries to manipulate and analyse data.
3. Acquire and integrate data from files, APIs and web pages, building preparation and transformation processes with data quality checks.
4. Conduct exploratory analysis and communicate findings through visualisations, reports and interactive dashboards.
5. Compare storage and querying models, selecting formats and technologies suited to data characteristics and usage requirements.
6. Explain parallel and distributed processing principles and apply large-scale processing strategies, assessing their costs and limitations.
7. Develop, document and defend practical data solutions, justify technical decisions, and critically verify results and assistance obtained from artificial intelligence tools.
Program
1. Introduction to data science and Big Data
- 1.1 Fundamental concepts, professional roles and the data project lifecycle.
- 1.2 The 5Vs: volume, velocity, variety, veracity and value.
- 1.3 Ethics, privacy, transparency, data quality and social impact.
- 1.4 Reproducibility, documentation, version control and critical use of AI.
2. Development environment and Python fundamentals
- 2.1 Jupyter, Python scripts and Visual Studio Code.
- 2.2 Isolated environments and dependency management: pip, conda and uv.
- 2.3 Docker, Docker Compose and, if time allows, an introduction to Dev Containers.
- 2.4 Python review: types, data structures, control flow, functions and modules.
- 2.5 Project organisation, Git, code checks and basic tests.
3. Data manipulation, analysis and visualisation
- 3.1 NumPy and array operations.
- 3.2 pandas: selection, transformation, aggregation and joins of tabular data.
- 3.3 Exploratory analysis, descriptive statistics, missing data and outliers.
- 3.4 Matplotlib, Seaborn and Plotly; chart selection and accurate communication.
- 3.5 Interactive dashboards with Streamlit and an introduction to alternatives such as Dash.
4. Data acquisition, integration and engineering
- 4.1 Files and formats: CSV, JSON, Excel and Parquet.
- 4.2 REST APIs: authentication, pagination, rate limits and error handling.
- 4.3 Web scraping with requests and Beautiful Soup; good collection practices.
- 4.4 Document data extraction and an introduction to OCR.
- 4.5 Cleaning, normalisation, integration and validation of data from multiple sources.
- 4.6 ETL and ELT pipelines, traceability and data updates.
5. Data storage and querying
- 5.1 Relational and NoSQL models: documents, key-value, column families and graphs.
- 5.2 SQL querying and analysis; examples with PostgreSQL and DuckDB.
- 5.3 Storage and querying with MongoDB and Redis.
- 5.4 Object storage, data warehouses, data lakes and an introduction to lakehouses.
- 5.5 Selection criteria: data model, access patterns, consistency, scalability and cost.
- 5.6 Comparative experiments with notebooks and local Docker services, where applicable.
6. Large-scale processing strategies
- 6.1 Memory and I/O constraints; assessing execution time and resource usage.
- 6.2 Chunking, vectorised operations and incremental processing.
- 6.3 Columnar formats, compression, partitioning and selective reads.
- 6.4 Parallelisation, lazy execution and an introduction to Dask.
7. Distributed processing
- 7.1 MapReduce and the fundamentals of distributing data and tasks.
- 7.2 Hadoop architecture: HDFS and YARN; the transition to Spark.
- 7.3 Apache Spark and PySpark: DataFrames, Spark SQL and an introduction to RDDs.
- 7.4 Transformations, actions, lazy execution, partitions and data transfers between tasks.
- 7.5 Comparing local, parallel and distributed solutions; measuring and interpreting performance.
- 7.6 Introduction to distributed machine learning with Spark MLlib and streaming processing (if time allows).
Evaluation Methodology
All assessment periods:
- 25%: Project I, data acquisition, preparation and exploratory analysis (5 points).
- 25%: Project II, data integration, storage and querying through an ETL/ELT pipeline (5 points).
- 25%: Project III, large-scale data processing with Spark and/or Dask (5 points).
- 25%: Theoretical examination (5 points).
Minimum grades:
- 50% in each project: 10/20, equivalent to 2.5/5.
- 35% in the examination: 7/20, equivalent to 1.75/5.
- A final grade of at least 9.5/20, with all component minima met.
The final grade is the sum of the four components graded out of 5, equivalent to the average of their grades out of 20.
Each project must be submitted and defended by the deadlines and at the times set during continuous assessment. Projects cannot be repeated or replaced by an examination. Failure to submit or defend any project excludes the student from course assessment in all periods of that academic year. Failure to meet any project minimum prevents a passing grade.
Project grades carry over across assessment periods within the same academic year. Only the theoretical examination may be repeated in applicable periods. Defences verify individual understanding and contributions, including AI use.
Bibliography
- Marr, B. (2022). Data Strategy: How to Profit from a World of Big Data, Analytics and the Internet of Things. USA: Kogan Page
- McKinney, W. (2017). Python for Data Analysis: Data Wrangling with Pandas, NumPy, and IPython. USA: O'Reilly
- Rioux, J. (2022). Data Analysis with Python and PySpark. USA: Manning
- Santos, M. e Costa, C. (2019). Big Data Concepts, Warehousing, and Analytics. . Lisboa: FCA
- Triguero, I. e Galar, M. (2023). Large-Scale Data Analytics with Python and Spark. UK: Cambridge University Press
Teaching Method
TP classes combine concepts, cases and practical tutorials. PL classes provide self-paced notebook and Docker experiments and project support. Independent work and critical use of AI support explanations, learning and solution validation.
Software used in class
Python, Jupyter, Visual Studio Code, pip, conda/Miniconda or micromamba, uv, Git, Docker and Docker Compose; NumPy, pandas, Matplotlib, Seaborn, Plotly, Streamlit, requests and Beautiful Soup; SQL, DuckDB, PostgreSQL, MongoDB, Redis, Dask and PySpark. Ruff and pytest support code quality and verification. Dev Containers, Dash and document extraction/OCR tools are used in supplementary examples.


















