Presented By O'Reilly and Cloudera
Make Data Work
5–7 May, 2015 • London, UK

Friction-free ETL: Automating data transformation with Impala

Marcel Kornacker (Cloudera)
10:55–11:35 Wednesday, 6/05/2015
Hadoop Platform
Location: King's Suite - Sandringham
Average rating: ****.
(4.08, 13 ratings)
Slides:   1-PPTX 

Prerequisite Knowledge

Basic knowledge of Hadoop ecosystem

Description

As data is ingested into Apache Hadoop at an increasing rate from a diverse range of data sources, it is becoming more and more important for users that new data be accessible for analysis as quickly as possible—because “data freshness” can have a direct impact on business results.

In the traditional ETL process, raw data is transformed from the source into a target schema, possibly requiring flattening and condensing, and then loaded into an MPP DBMS. However, this approach has multiple drawbacks that make it unsuitable for real-time, “at-source” analytics—for example, the “ETL lag” reduces data freshness, and the inherent complexity of the process makes it costly to deploy and maintain, and reduces the speed at which new analytic applications can be introduced.

In this talk, attendees will learn about Impala’s approach to on-the-fly, automatic data transformation, which in conjunction with the ability to handle nested structures such as JSON and XML documents, addresses the needs of at-source analytics—including direct querying of your input schema, immediate querying of data as it lands in HDFS, and high performance on par with specialized engines. This performance level is attained in spite of the most challenging and diverse input formats, which are addressed through an automated background conversion process into Parquet, the high-performance, open source columnar format that has been widely adopted across the Hadoop ecosystem.

In this talk, attendees will learn about Impala’s upcoming features that will enable at-source analytics: support for nested structures such as JSON and XML documents, which allows direct querying of the source schema; automated background file format conversion into Parquet, the high-performance, open source columnar format that has been widely adopted across the Hadoop ecosystem; and automated creation of declaratively-specified derived data for simplified data cleansing and transformation.

Photo of Marcel Kornacker

Marcel Kornacker

Cloudera

Marcel Kornacker is the architect and tech lead at Cloudera for Impala. Prior to Cloudera, Marcel worked at Google on several ad-serving and storage infrastructure projects. He eventually became the tech lead for the distributed query engine component of Google’s F1 project. He holds a PhD in databases from UC Berkeley.