教學大綱與進度
課程基本資料:
學年期
課號
課程名稱
階段
學分
時數
修
教師
班級
人
撤
備註
112-1
325089
大數據技術與管理
1
3.0
3
★
王正豪
王怡鈞
AI雙聯專班
13
0
碩一修,9/24,10/15,11/12,12/3,12/9,12/23
教學大綱與進度:
教師姓名
王怡鈞
Email
bnwntut8@ntut.edu.tw
最後更新時間
2023-09-20 14:06:35
課程大綱
This is an applied science course that examines how the application of big data can enhance technical and functional capabilities and how the big data processing efficiency can be facilitated and monitored to achieve well-defined objectives. As it is well known that an optimized result made by Artificial Intelligence (AI) could not be significantly achieved without big data. Nowadays, there are a wide variety of data type (mostly non-structured ones) in different applications, and we are generating much more data and faster than ever before. Thus, big data applications are impacting our daily life. However, there are problems using existing techniques to accommodate most of our interests. For example, accessing the data from disk in distributed systems and getting data to the processors are among the typical bottlenecks. This course is designed to not only introduce Big Data application and trend with Hadoop, but also its processing technology and ecosystems which are capable of eliminating the bottlenecks by its unique methods for storing and processing data per common application scenario. In addition, the course provides the platforms for students’ paper reading, group/squad discussion, co-work assignment, case studies and presentation. The course will introduce Apache Hadoop core technologies and administration with its major components including: Big Data (BD) in demand and its application trend, Core technologies of Hadoop for BD management, Hadoop Administration and management tool, Hadoop cluster deployment, YARN applications plus MapReduce, Hadoop cluster hardware and software, Introduction of Hadoop ecosystem projects, HDFS configuration for high availability (HA), Hadoop security, Cluster resource scheduling , hardware configuration, industrial case studies, forum/project discussion, etc.. In addition, due to the major bottleneck of “shuffle and sort” in MapReduce that affects the performance of Hadoop, we will also introduce the memory-based distributed framework Spark. The topics include: Spark API, RDDs, Data Frames. The most common data mining applications including classification and clustering will be introduced to show the differences between Hadoop and Spark. 由於雲端運算的快速發展,大量資料收集與儲存技術已成瓶頸,大數據技術與管理提供一套具延展性與容錯能力的分散儲存與計算架構, 能有效處理大量複雜數據資料的儲存與運算。 本課程除介紹現有的大數據生態與趨勢,並針對目前軟硬體架構的處理瓶頸,提供系統組態與部署、服務層級協定、監控與管理等系統性規劃的解決方案,並導入跨領域的工程管理科學知識。課程簡扼論述大數據技術與管理的重要,並在Hadoop生態系統中,以Cloudera Hadoop為應用範例,實作說明其重要性。 課程內容包括:大數據的概念、應用生態與趨勢、Hadoop硬體規格與基礎架構規劃、叢集組態與部署、叢集服務層級協定、叢集監控與管理、相關生態技術、業界案例說明與專案研習討論。此外,由於MapReduce運算過程可能耗費太多磁碟I/O,為了能進一步提升大數據處理的效率,課程亦將介紹基於記憶體運算的Spark架構,包括:Spark API、RDD、Data Frame等; 並且以常見的資料探勘應用範例,包括: 分類、分群等,來說明Spark叢集的運作、Spark與Hadoop的差異。
課程進度
(1-16週)
(王怡鈞 老師) - Session-1 • Unit 1 Big Data Demand and Trend • Unit 2 Hadoop Distributed File System (HDFS) (王怡鈞 老師) - Session-2 • Unit 3 MapReduce & Yet Another Resource Negotiator (YARN) • Unit 4 YARN Scheduler and Applications (王怡鈞 老師) - Session-3 • Unit 5 Hadoop Cluster Planning and Deployment • Unit 6 Hadoop Security and Administration (王怡鈞 老師) - Session-4 • Unit 7 Big Data Industrial Case Studies • Unit 8 Project Forum/Discussion with Teamwork
評量方式與標準
Term Assignments, Study assessment, 67% (王怡鈞 老師) 1 Team Assignment(s): Reports, Presentations, Q&A, Commenting, 40% weight 2 Exam, 50% weight 3 Individual Participations in class activities, 10% weight
使用教材、參考書目或其他
【遵守智慧財產權觀念,請使用正版教科書,不得使用非法影印教科書】
使用外文原文書:是
In-class handout and research papers downloaded from internet. Research Paper Readings – (Google’s issued papers) 1. The Google File System; https://static.googleusercontent.com/media/research.google.com/zh-TW//archive/gfs-sosp2003.pdf 2. MapReduce: Simplified Data Processing on Large Clusters; https://static.googleusercontent.com/media/research.google.com/zh-TW//archive/mapreduce-osdi04.pdf 3. Big Data associated, To-Be-Advised in class Additional Reading Reference 1. Hadoop: The Definitive Guide, Fourth Edition, April 2015, Published by O’Reilly Media, Inc., Editor: Tom White. 2. Hadoop Essentials, First Edition, April 2015, Published by O’Reilly Media, Inc., Editor: Shiva Achari. 3. Hadoop Operation, First Edition, September 2012, Published by O’Reilly Media, Inc., Editor: Eric Sammer. 4. 大數據大時代 - 新一代儲存技術及實作,作者: 查偉,發行所:佳魁資訊, 出版日期: 2017, 3月. 5. Yusuf Aytas, Designing Big Data Platform, How to Use, Deploy and Maintain Big Data Systems, Wiley, February 2021. 6. Jimmy Lin and Chris Dyer, Data-Intensive Text Processing with MapReduce, Morgan & Claypool Publishers, 2010. 7. Holden Karau, Andy Konwinski, Patrick Wendell, Matei Zaharia, Learning Spark: Lightning-Fast Big Data Analysis, O'Reilly Media, January 2015.
課程諮詢管道
School Email Address: bnwntut8@ntut.edu.tw
LINE Group: https://line.me/R/ti/g/NjjtWakuKa
延伸教學與資源
課程對應SDGs指標
課程是否導入AI
備註
SDG 4: Quality Education SDG 8: Decent Work and Economic Growth SDG 9: Industry, Innovaton and Infrastructure SDG 11: Sustainable Cities and Communities