教學大綱與進度
課程基本資料:
學年期
課號
課程名稱
階段
學分
時數
修
教師
班級
人
撤
備註
111-2
313504
大數據技術與管理
1
3.0
3
★
王怡鈞
電機所
人工智慧學位學程
創新AI學位學程
22
0
日職電機所和AI學程合開
教學大綱與進度:
教師姓名
王怡鈞
Email
bnwntut8@ntut.edu.tw
最後更新時間
2023-01-03 12:21:40
課程大綱
This is a science applied course that examines how the application of big data is made for our human to enhance technical and functional capabilities and how the big data processing efficiency can be facilitated and monitored to achieve well-defined objectives. As it is well known that an optimized result made by Artificial Intelligence (AI) could not be significantly achieved without big data feedings. Nowadays, there are diversity of data type (mostly of non-structured ones) spreading on the air and pipelines, furthermore we are generating much more data and faster than ever before, and the application of big data does impact our daily life. Although we can store and process the data, there are problems using existing techniques to accommodate our most interests. For example, both accessing the data from disk in distributed systems and getting data to the processors are of the typical bottlenecks. This course is designed to not only introduce students to Hadoop Big Data application and trend, but also its processing technology and ecosystems which are capable of eliminating the bottlenecks by its unique methods for storing and processing data per common application scenario. In addition, the course provides the platforms for students’ paper reading, group discussion, co-work assignment, case studies and presentation. The course will introduce Hadoop core technologies and administration with its major listed below: Big Data in Demand and Its Application Trend Core Technologies of Hadoop for Big Data Management Management tool to simplify Hadoop Administration Hadoop Cluster Deployment YARN Applications including MapReduce Hadoop Cluster Hardware and Software Deployment of Hadoop Ecosystem Projects: Sqoop, Flume, Hive and Impala HDFS Configuration for High Availability Hadoop Security Cluster Resource Scheduling and Configuration Cluster Monitoring and Management Industrial Big Data and Ecosystems Development and Case Studies 由於雲端運算的快速發展,大量資料收集與儲存技術已成瓶頸,大數據技術與管理提供一套具延展性與容錯能力的分散儲存與計算架構, 能有效處理大量複雜數據資料的儲存與運算。 本課程除介紹現有的大數據生態與趨勢,並針對目前軟硬體架構的處理瓶頸,提供系統組態與部署、服務層級協定、監控與管理等系統性規劃的解決方案,並導入跨領域的工程管理科學知識。課程簡扼論述大數據技術與管理的重要,並在Hadoop生態系統中,以Cloudera Hadoop為應用範例,實作說明其重要性。 課程內容包括:大數據的概念、應用生態與趨勢、Hadoop硬體規格與基礎架構規劃、叢集組態與部署、叢集服務層級協定、叢集監控與管理、相關生態技術。
課程進度
(1-16週)
(Syllabus, preliminary) Tuesday. 18:30-21:10 (ABC) Week# Subject to change per needs and circumstances 1 2/21 Syllabus & organization of all mission cases from the class 2 2/28 Public Holiday 3 3/07 Big Data Application and Trend 4 3/14 Core Technologies for Big Data Management 5 3/21 Hadoop Cluster Deployment 6 3/28 Hadoop Distributed File System (HDFS) 7 4/04 Public Holiday 8 4/11 MapReduce 9 4/18 YARN Applications 10 4/25 Planning Hadoop Cluster Hardware and Software 11 5/02 Hadoop I/O 12 5/09 Hadoop High Available, Failover, Fensing and Federation 13 5/16 Hadoop Ecosystem Project Presentation and Q&A -1 14 5/23 Hadoop Ecosystem Project Presentation and Q&A -2 15 5/30 Hadoop Ecosystem Project Presentation and Q&A -3 16 6/06 Resource & Scheduler Management 17 6/13 Hodoop Security 18 6/20 Final-Exam
評量方式與標準
1. Assignment including Presentation, Comments and Q&A:50% 2. Final Exam: 40% 3. Participations in class activities: 10%
使用教材、參考書目或其他
【遵守智慧財產權觀念,請使用正版教科書,不得使用非法影印教科書】
使用外文原文書:是
1. Hadoop: The Definitive Guide, Fourth Edition, April 2015, Published by O’Reilly Media, Inc., Editor: Tom White. 2. Hadoop Essentials, First Edition, April 2015, Published by O’Reilly Media, Inc., Editor: Shiva Achari. 3. Hadoop Operation, First Edition, September 2012, Published by O’Reilly Media, Inc., Editor: Eric Sammer. 4. 大數據大時代 - 新一代儲存技術及實作,作者: 查偉,發行所:佳魁資訊, 出版日期: 2017, 3月. 5. Yusuf Aytas, Designing Big Data Platform, How to Use, Deploy and Maintain Big Data Systems, Wiley, February 2021. 6. Jimmy Lin and Chris Dyer, Data-Intensive Text Processing with MapReduce, Morgan & Claypool Publishers, 2010. 7. Holden Karau, Andy Konwinski, Patrick Wendell, Matei Zaharia, Learning Spark: Lightning-Fast Big Data Analysis, O'Reilly Media, January 2015. Additional Reading References: 8. The Google File System; https://static.googleusercontent.com/media/research.google.com/zh-TW//archive/gfs-sosp2003.pdf 9. MapReduce: Simplified Data Processing on Large Clusters; https://static.googleusercontent.com/media/research.google.com/zh-TW//archive/mapreduce-osdi04.pdf 10. Big Data associated, To-Be-Advised in class
課程諮詢管道
Pls
Join LINE Group: To be advised(TBA)
Questions to email: bnwntut8@ntut.edu.tw
延伸教學與資源
課程對應SDGs指標
課程是否導入AI
備註
If NTUT anounces for remote class due to COVID-19, then it goes for Google Meet: https://meet.google.com/rkg-kbgv-odo Pls join LINE Group: To be advised (TBA) Questions to email: bnwntut8@ntut.edu.tw