DeepDive is a trained system that uses machine learning to cope with various forms of noise and imprecision. DeepDive is designed to make it easy for users who do not have machine-learning expertise to train the system through low-level feedback via the MindTagger interface and discover rich, structured domain knowledge via rules. Mike Cafarella offers an introduction to DeepDive, exploring the key technical innovations that enable DeepDive to produce statistical inference at massive scale.
Mike Cafarella is one of the cofounders of the Apache Hadoop and Nutch open source projects. Mike is also an assistant professor of computer science and engineering at the University of Michigan. His research interests include databases, information extraction, data integration, and data mining. Recently, he cofounded Lattice Data (http://lattice.io), a company that aims to transform “dark data,” such as unstructured text documents and reports, into high quality structured databases.
©2016, O'Reilly Media, Inc. • (800) 889-8969 or (707) 827-7019 • Monday-Friday 7:30am-5pm PT • All trademarks and registered trademarks appearing on oreilly.com are the property of their respective owners. • firstname.lastname@example.org
Apache Hadoop, Hadoop, Apache Spark, Spark, and Apache are either registered trademarks or trademarks of the Apache Software Foundation in the United States and/or other countries, and are used with permission. The Apache Software Foundation has no affiliation with and does not endorse, or review the materials provided at this event, which is managed by O'Reilly Media and/or Cloudera.