InfoQ Homepage Hadoop Content on InfoQ

News

RSS Feed

Newer Older

AI, ML & Data Engineering

Using Hunk+Hadoop as a Backend for Splunk

Splunk can now store archived indexes on Hadoop. At the cost of performance, this offers a 75% reduction in storage costs without losing the ability to search the data. And with the new adapters, Hadoop tools such as Hive and Pig can process the Splunk-formatted data.

Jonathan Allen
on Sep 22, 2015
AI, ML & Data Engineering

Splunk .conf 2015 Keynote

Splunk opened their big data conference with an emphasis on “making machine data accessible, usable, and valuable to everyone”. This is a shift from their original focus: indexing arbitrary big data sources. Reasonably happy with their ability to process data, they want to ensure that developers, IT staff, and normal people have a way to actually use all of the data their company is collecting.

Jonathan Allen
on Sep 22, 2015
Parquet Becomes Top-Level Apache Project

Apache Parquet, the open-source columnar storage format for Hadoop, recently graduated from the Apache Software Foundation Incubator and became a top-level project. Initially created by Cloudera and Twitter in 2012 to speed up analytical processing, Parquet is now openly available for Apache Spark, Apache Hive, Apache Pig, Impala, native MapReduce, and other key components of the Hadoop ecosystem.

Jérôme Serrano
on Jun 11, 2015
MemSQL 4 Database Supports Community Edition, Geospatial Intelligence and Spark Integration

Latest version of MemSQL, in-memory database with support for transactions and analytics, includes a new Community Edition for free use by organizations. MemSQL 4, released last week, also supports integration with Apache Spark, Hadoop Distributed File System (HDFS), and Amazon S3.

Srini Penchikala
on May 30, 2015
Hortonworks, IBM and Pivotal to Support Open Data Platform in Their Big Data Solutions

Big data vendors Hortonworks, IBM, and Pivotal recently announced that their Hadoop based platform products will use the common Open Data Platform (ODP). They made the announcement at the recent HadoopSummit Europe Conference of the open platform which includes Apache Hadoop 2.6 (HDFS, YARN, and MapReduce) and Apache Ambari software.

Srini Penchikala
on Apr 24, 2015
Apache HBase Hits 1.0

After three developer previews, six release candidates and over 1500 closed tickets the Apache foundation has announced version 1.0 of Apache HBase, a NoSQL database in the Hadoop ecosystem. After more than 7 years of active development, the team behind HBase felt that the project had matured and stabilized enough to warrant a 1.0 version.

Benjamin Darfler
on Apr 07, 2015
Spring XD 1.1: Simplifying Big Data like Spring Did for Java EE

Pivotal recently released Spring XD 1.1 GA with new features including stream processing with Reactor, RxJava, Spark Streaming and Python. Additionally support for Kafka, batching and compression with RabbitMQ, and support for container group management when running on YARN are now featured.

Matt Raible
on Mar 05, 2015
Google Open Sources MapReduce Framework for C to Run Native Code in Hadoop

Google announced last week the release of open source MapReduce framework for C, called MR4C, that allows developers to run native code in Hadoop framework. MR4C framework brings together the performance and flexibility of natively developed algorithms with the scalability and throughput provided by Hadoop execution framework.

Srini Penchikala
on Feb 25, 2015
Project Pachyderm Aims to Build a "Modern" Hadoop on Docker

Project Pachyderm Aims to Build "Modern" Hadoop using Docker and CoreOS.

Matt Kapilevich
on Feb 17, 2015
Apache Hive 1.0 Released, HiveServer2 Becomes Main Engine, Stable API Defined

Apache Hive has released version 1.0 of their project on February 6th, 2015. Originally planned as version 0.14.1, the community voted to change the version numbering to 1.0.0 to reflect the amount of maturity the project has reached.

Mikio Braun
on Feb 11, 2015
Splice Machine Version 1.0 Supports Integration with Hadoop and Analytic Window Functions

Splice Machine version 1.0 supports analytic window functions and integration with Hadoop ecosystem. Splice Machine team recently released their Hadoop based RDBMS data management solution that can be used for transactional workloads on Hadoop.

Srini Penchikala
on Dec 18, 2014
LinkedIn Open Sources Cubert With an Eye To Big Data Analytics

LinkedIn recently open sourced Cubert, its High Performance Computation Engine for Complex Big Data Analytics. Cubert is a framework written for analysts and data scientists in mind.Developed completely in Java and expressed as a scripting language, Cubert is designed for complex joins and aggregations that frequently arise in the reporting world.

Alex Giamas
on Dec 17, 2014
Gobblin, LinkedIn's Unified Data Ingestion Platform

At the 2014 QCon San Francisco conference, LinkedIn's Lin Qiao gave a talk on their Gobblin project (also summarized in a blog post) that is a unified data ingestion system for their internal and external data sources.

Mikio Braun
on Dec 15, 2014
Stripe Open Sources Tools For Apache Hadoop

Stripe, the internet payments infrastructure company recently announced open sourcing a set of internally developed tools based on Apache Hadoop.Timberlake, Brushfire, Sequins and Herringbone all contribute to enriching the available tools for building an Apache Hadoop stack.

Alex Giamas
on Dec 09, 2014
Spark Sets New Record in Sort Performance

Databricks has recently announced a new record in the Daytona GraySort contest using the Spark processing engine. The Daytona GraySort contest is a 3rd party benchmark measuring how fast a system can sort 100 Terabytes of data. Databricks posted a throughput of 4.27 TB/min over a cluster of 206 machines for their official run.

Benjamin Darfler
on Nov 26, 2014

Newer News

Older News

InfoQ Software Architects' Newsletter

News