Presented By O'Reilly and Cloudera
Make Data Work
31 May–1 June 2016: Training
1 June–3 June 2016: Conference
London, UK

HopsWorks: Multitenant Hadoop as a service

Jim Dowling (Swedish ICT - SICS)
16:35–17:15 Friday, 3/06/2016
Security
Location: Capital Suite 15/16 Level: Intermediate
Average rating: *****
(5.00, 1 rating)

Prerequisite knowledge

Attendees should have experience with data science and writing programs in Python, Scala, or Java.

Description

Jim Dowling describes how HopsWorks enables organizations to securely share a single Hadoop cluster using projects and a new metadata layer that enables protection domains while still allowing data sharing.

HopsWorks is a frontend to Hadoop that provides a new model for multitenancy in Hadoop, based around projects. A project is like a GitHub project: the owner of the project manages membership, and users can be one of two roles in the project—data scientists, who can just run programs, and data owners, who can also curate, import, and export data. Users can’t copy data between projects or run programs that process data from different projects, even if the user is a member of multiple projects. That is, we implement multitenancy with dynamic roles, where the user’s role is based on the currently active project. (Users can still share datasets between projects, however.)

HopsWorks also supports Apache Zeppelin with access control, free-text search for files in HDFS using Elasticsearch, and a metadata designer tool that enables users to curate files and directories in HDFS. HopsWorks has been enabled by migrating all metadata in HDFS and YARN into an open source, shared nothing, in-memory, distributed database called NDB. HopsWorks is open source and licensed as Apache v2, with database connectors licensed as GPL v2. From late January 2016, HopsWorks will be provided as software as a service for researchers and companies in Sweden by the Swedish ICT SICS Data Center.

Photo of Jim Dowling

Jim Dowling

Swedish ICT - SICS

Jim Dowling is a senior researcher at SICS – Swedish ICT and an associate professor at KTH Stockholm. Jim is a researcher in the areas of distributed systems, machine learning, and large-scale computer systems. He also worked as a senior consultant for MySQL AB. Jim is lead architect of Hadoop Open Platform as a Service, a more scalable and highly available distribution of the Hadoop.