Pentaho Data Science Pack Operationalizes Use of R and Weka
According to the O'Reilly Data Scientist Salary Survey, R is the most-used tool for data scientists, while Weka is a widely used and popular open source collection of machine learning algorithms. Today, Pentaho Corporation announced the Data Science Pack, a toolkit to operationalize these two commonly used technologies. Through Pentaho Data Integration (PDI), data scientists can offload the drudgery of the data flow process with analytic components for R and Weka, so organizations can spend more time on strategic advanced and predictive analytics to achieve a more complete view into customer behavior.
“There was a gap in the market until now and people like myself were piecing together solutions to help with the data preparation, cleansing and orchestration of analytic data sets. The Pentaho Data Science Pack fills that gap to operationalize the data integration process for advanced and predictive analytics,” said Ken Krooner, President at ESRG. “Having embedded Pentaho for over seven years to provide remote and onboard analytics for maritime fleets and ships and several years experience with different data tools, Pentaho Data Integration is critical to my team. Using Weka with PDI, we are now helping clients blend a 360-degree view of all equipment data sources to enable early prediction of potential machinery failure.
According to the Ventana Research Big Data Analytics Benchmark Research, the top two time-consuming big data tasks are solving data quality and consistency issues (46%) and preparing data for integration (52%). Pentaho customer Paytronix manages marketing and loyalty programs in the restaurant sector and their team of data scientists uses R with Pentaho and Hadoop to analyze data to help their customers predict fraud and buying behavior. Saad Khalid, Data Insights Product Manager at Paytronix explains, “Data preparation is an essential, but tedious process. Pentaho Data Integration with R has allowed Paytronix to deliver analytics and insights to our clients much faster. What used to take a few weeks now takes a few minutes.”
“Having built blueprints for the four most popular big data use cases, Pentaho is at the forefront of solving data integration challenges, and we know advanced and predictive analytics are core ingredients for success,” said Christopher Dziekan, EVP and Chief Product Officer at Pentaho. “The highest value of data comes from having foresight blended with hindsight to drive insight and action. The Pentaho Data Science Pack allows organizations to apply their deep domain expertise and improve their customer analytics and predictions.”
The Data Science Pack improves productivity by executing advanced descriptive statistics and machine learning algorithms “at scale” inside of data flow transformations. Included in the pack are:
- The R Script Executor step enables the more than 5,500 packages in the Comprehensive R Archive Network (CRAN) repository to be used within PDI transformations
- The Weka Forecasting step uses machine learning techniques to generate forward-looking time-series data sets based on historical observations
- The Weka Scoring step executes machine learning models to calculate and append probability values to incoming records
The Data Science Pack will be available from Pentaho early this summer.
About Pentaho Corporation
Pentaho is delivering the future of business analytics. Pentaho's open source heritage drives our continued innovation in a modern, integrated, embeddable platform built for the future of analytics, including diverse and big data requirements. Powerful business analytics are made easy with Pentaho's cost-effective suite for data access, visualization, integration, analysis and mining. For a free evaluation, download Pentaho Business Analytics at www.pentaho.com/get-started.
Pentaho Media Contact
US & Worldwide
Director of Corporate Communications