Apache Hive Essentials-Packt Publishing(2015).pdf

发布时间：2022-06-24 发布人：admin 分类：说明书资料大小：2.17M 资料格式：pdf 举报版权申诉

ebfbf9b3-04bd-4a2a-bd77-67276465fe46.pdf-第1页.png

第1页 / 共208页

ebfbf9b3-04bd-4a2a-bd77-67276465fe46.pdf-第2页.png

第2页 / 共208页

ebfbf9b3-04bd-4a2a-bd77-67276465fe46.pdf-第3页.png

第3页 / 共208页

ebfbf9b3-04bd-4a2a-bd77-67276465fe46.pdf-第4页.png

第4页 / 共208页

ebfbf9b3-04bd-4a2a-bd77-67276465fe46.pdf-第5页.png

第5页 / 共208页

ebfbf9b3-04bd-4a2a-bd77-67276465fe46.pdf-第6页.png

第6页 / 共208页

ebfbf9b3-04bd-4a2a-bd77-67276465fe46.pdf-第7页.png

第7页 / 共208页

ebfbf9b3-04bd-4a2a-bd77-67276465fe46.pdf-第8页.png

第8页 / 共208页

Cover

Credits

About the Author

About the Reviewers

www.PacktPub.com

Table of Contents

Preface

Chapter 1: Overview of Big Data and Hive

A short history

Introducing big data

Relational and NoSQL database versus Hadoop

Batch, real-time, and stream processing

Overview of the Hadoop ecosystem

Hive overview

Summary

Chapter 2: Setting Up the Hive Environment

Installing Hive from Apache

Installing Hive from vendor packages

Starting Hive in the cloud

Using the Hive command line and Beeline

The Hive-integrated development environment

Summary

Chapter 3: Data Definition and Description

Understanding Hive data types

Data type conversions

Hive Data Definition Language

Hive database

Hive internal and external tables

Hive partitions

Hive buckets

Hive views

Summary

Chapter 4: Data Selection and Scope

The SELECT statement

The INNER JOIN statement

The OUTER JOIN and CROSS JOIN statements

Special JOIN – MAPJOIN

Set operation – UNION ALL

Summary

Chapter 5: Data Manipulation

Data exchange – LOAD

Data exchange – INSERT

Data exchange – EXPORT and IMPORT

ORDER and SORT

Operators and functions

Transactions

Summary

Chapter 6: Data Aggregation and Sampling

Basic aggregation – GROUP BY

Advanced aggregation – GROUPING SETS

Advanced aggregation – ROLLUP and CUBE

Aggregation condition – HAVING

Analytic functions

Sampling

Summary

Chapter 7: Performance Considerations

Performance utilities

The EXPLAIN statement

The ANALYZE statement

Design optimization

Partition tables

Bucket tables

Index

Data file optimization

File format

Compression

Storage optimization

Job and query optimization

Local mode

JVM reuse

Parallel execution

Join optimization

Common join

Map join

Bucket map join

Sort merge bucket (SMB) join

Sort merge bucket map (SMBM) join

Skew join

Summary

Chapter 8: Extensibility Considerations

User-defined functions

The UDF code template

The UDAF code template

The UDTF code template

Development and deployment

Streaming

SerDe

Summary

Chapter 9: Security Considerations

Authentication

Metastore server authentication

HiveServer2 authentication

Authorization

Legacy mode

Storage-based mode

SQL standard-based mode

Encryption

Summary

Chapter 10: Working with Other Tools

JDBC/ODBC connector

HBase

Hue

HCatalog

ZooKeeper

Oozie

Hive roadmap

Summary

Index

Apache Hive Essentials Immerse yourself on a fantastic journey to discover the attributes of big data by using Hive Dayong Du BIRMINGHAM - MUMBAI

Apache Hive Essentials Copyright © 2015 Packt Publishing All rights reserved. No part of this book may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, without the prior written permission of the publisher, except in the case of brief quotations embedded in critical articles or reviews. Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However, the information contained in this book is sold without warranty, either express or implied. Neither the author, nor Packt Publishing, and its dealers and distributors will be held liable for any damages caused or alleged to be caused directly or indirectly by this book. Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals. However, Packt Publishing cannot guarantee the accuracy of this information. First published: February 2015 Production reference: 1210215 Published by Packt Publishing Ltd. Livery Place 35 Livery Street Birmingham B3 2PB, UK. ISBN 978-1-78355-857-5 www.packtpub.com

Credits Project Coordinator Neha Bhatnagar Proofreaders Paul Hindle Jonathan Todd Indexer Monica Ajmera Mehta Production Coordinator Aparna Bhagat Cover Work Aparna Bhagat Author Dayong Du Reviewers Puneetha B M Hamzeh Khazaei Nitin Pradeep Kumar Balaswamy Vaddeman Commissioning Editor Ashwin Nair Acquisition Editor Shaon Basu Content Development Editor Merwyn D'souza Technical Editor Taabish Khan Copy Editors Sameen Siddiqui Laxmi Subramanian

About the Author Dayong Du is a big data practitioner, leader, and developer with expertise in technology consulting, designing, and implementing enterprise big data solutions. With more than 10 years of experience in enterprise data warehouse, business intelligence, and big data and analytics, he has provided his data intelligence expertise in various industries, such as media, travel, telecommunications, and so on. He is currently working with QuickPlay Media in Toronto, Canada, to build enterprise big data intelligence reporting for online media services and content providers. He has a master's degree in computer science from Dalhousie University, and he holds the Cloudera Certified Developer for Apache Hadoop certification. I would like to sincerely thank my wife, Joice, and daughter, Elaine, for their sacrifices and encouragement during this journey. Also, I would like to thank my parents for their support during the time of writing this book. I would also like to thank everyone at Packt Publishing and the technical reviewers for their valuable help, guidance, and feedback on my book.

About the Reviewers Puneetha B M is a software engineer, data enthusiast, and technical blogger. Her research interests include big data, cloud computing, machine learning, and NoSQL databases. She is also a professional software engineer with more than 2 years of working experience. She holds a master's degree in computer applications from P.E.S. Institute of Technology. Other than programming, she enjoys painting and listening to music. You can learn more from her blog (http://blog.puneethabm. in/) and LinkedIn profile (https://www.linkedin.com/in/puneethabm). I owe a great deal to Prof. Dr. Ram Rustagi for being a role model in my life and for his zealous inspiration. I would like to thank my brother, Nischith B.M., for supporting me in everything I do. I would also like to thank Packt Publishing and its staff for providing the opportunity to contribute to this book. Hamzeh Khazaei is a postdoctoral research scientist at IBM Canada Research and Development Centre. He received his PhD degree in computer science from University of Manitoba, Winnipeg, Manitoba, Canada (2009–2012). Earlier, he received both his BSc and MSc degrees in computer science from Amirkabir University of Technology, Tehran, Iran (2000–2008). He is also a sessional instructor in the Computer Science department at Ryerson University (http://scs.ryerson.ca/~hkhazaei). He teaches software engineering to fourth year undergraduate students. His research area includes big data analytics, cloud computing infrastructure, analytics as a service, and modeling of computing systems. I would like to thank my dear wife for her perpetual support in all my endeavors.

Nitin Pradeep Kumar is a passionate developer with extensive experience and oodles of interest in emerging technologies such as the cloud and mobile. He is currently a cloud quality engineer at Appcelerator, a leading Silicon Valley-based start-up that provides an MBaaS platform purpose-built for mobile and cloud development. Before this stint, he studied at the National University of Singapore toward a master's degree in knowledge engineering, which involves building intelligent systems using cutting-edge artificial intelligence and data-mining techniques. He enjoys the start-up environment and has worked with technologies such as Hadoop, Hive, and data warehousing. He lives in Singapore and spends his spare cycles playing retro PC games on his mobile and learning Muay Thai. I would like to thank my family, friends, and my wonderful brother, Nivin, for supporting me in all my endeavors. Balaswamy Vaddeman is a Hadoop hackathon winner for Hyderabad in 2013. He is one of the top contributors on the Hive tag at http://www.stackoverflow. com. He is a big data professional with 3 years of experience. He is well known for training people on big data/Hadoop. So far, he has delivered six big data projects. He is a Java/J2EE expert with 8 years of IT experience and 5 years of RDBMS experience. He is an automation expert on Unix-based systems using Shell scripting. He has experience in setting up teams and bringing them up to speed on big data projects. He is an active participant in Hadoop/big data forums. I would like to thank my wife, Radha, my son, Pandu, and my daughter, Bubly, for their cooperation in completing this book.

www.PacktPub.com Support files, eBooks, discount offers, and more For support files and downloads related to your book, please visit www.PacktPub.com. Did you know that Packt offers eBook versions of every book published, with PDF and ePub files available? You can upgrade to the eBook version at www.PacktPub. com and as a print book customer, you are entitled to a discount on the eBook copy. Get in touch with us at service@packtpub.com for more details. At www.PacktPub.com, you can also read a collection of free technical articles, sign up for a range of free newsletters and receive exclusive discounts and offers on Packt books and eBooks. TM https://www2.packtpub.com/books/subscription/packtlib Do you need instant solutions to your IT questions? PacktLib is Packt's online digital book library. Here, you can search, access, and read Packt's entire library of books. Why subscribe? • Fully searchable across every book published by Packt • Copy and paste, print, and bookmark content • On demand and accessible via a web browser Free access for Packt account holders If you have an account with Packt at www.PacktPub.com, you can use this to access PacktLib today and view 9 entirely free books. Simply use your login credentials for immediate access.

分享到：

赞收藏

资料库

Apache Hive Essentials-Packt Publishing(2015).pdf

相关推荐

数据库

热门标签

最新资料