Error message
User error : Failed to connect to memcache server: druportbe01:11211 in dmemcache_object() (line 415 of /production/drupal/dim_prod/drupal/d7cl4/prod/unich/releases/7/web/sites/all/modules/contrib/memcache/dmemcache.inc ).
Course Sheet Academic Year of enrolment:
Professor and Collaborators:
Hours of classroom activity:
Prerequisites:
The knowledge of Basic Statistics is requested. Moreover, the knowledge of basic programming is suggested.
Objectives
Contents - Introduction to the Big Data phenomenon
- Big data methods
- Data mining
- Lab & tools
Extended Syllabus Introduction
• Introduction to the Big Data
Big data methods
• Programming frameworks: MapReduce/Hadoop, Spark
Data mining
• Association Analysis
• Clustering
• Graph Analytics (centrality measures, scale-free/Power-law graphs, small world phenomenon, uncertain graphs)
• Similarity and diversity search
Lab & tools
• tools and methodologies for collecting, processing, visualizing and analyzing large amounts of data (Big Data).
o extract unstructured data from web (import.io, kimono, etc.)
o explore and present static data (RAWGraphs, Gephi, illustrator, etc.)
o explore and build interactive data visualizations (Tableau Public, Carto)
Recommended Bibliography Course slides.
Further readings:
Leskovec, Jure, Anand Rajaraman, and Jeffrey David Ullman.
Mining of massive datasets.
Cambridge University Press, 2014.
Online available for free: http://www.mmds.org/
Teaching Methods Lectures.
Practice and exercises in the computer lab.
Evaluation methods Verification of learning:
Knowledge and understanding
The verification of the learning outcomes will be carried out through a written and oral examination (the latter being optional or potentially required by the teacher). The score of the exam is assigned by a mark expressed in 30ths and is based on both the written and oral examinations.
Applying knowledge and understanding
During the exam, students' ability to apply the knowledge given in the course is verified. In particular, student should be able to extract and manipulate data from the web, from files and from databases, including those of big size.