Welcome!

@BigDataExpo Authors: Elizabeth White, Pat Romanski, Liz McMillan, Yeshim Deniz, William Schmarzo

Related Topics: @BigDataExpo, Java IoT, Open Source Cloud, Containers Expo Blog, Agile Computing, @CloudExpo

@BigDataExpo: Article

Big Data: So What! That’s Why You Virtualize

How data virtualization enables Big Data volume, variety, velocity and value

Big Data!  Yes it's BIG!

The volume is BIG!  The variety is BIG!  The velocity is BIG!

And hopefully the business value is BIG!

New Opportunities Bring New Ways to Leverage Proven Technology
There is no shortage of media articles, analyst reports, tradeshows, blogs and other source of Big Data technology insight and advice.

But it strikes me that in our search to be on the leading edge, we may be overlooking some great existing technology.

In fact, some technology, for example data virtualization, is even more useful in a Big Data world.

What Is Data Virtualization?
Data virtualization is an agile data integration approach organizations use to gain more insight from their data.  This includes traditional sources such as transaction systems, data warehouses and more as well as new sources such the cloud and Big Data.

Unlike data consolidation or data replication, data virtualization integrates these diverse data types without costly extra copies and additional data management complexity.  Seriously, if the data is already big, why make it even bigger by copying and storing it again and again?

With data virtualization, you respond faster to ever changing analytics and BI needs, fast-track your data management evolution and save 50-75% over data replication and consolidation.   In other words, you deliver value, the most important V but often not listed with the 3 Vs of Big Data (Volume, Velocity & Variety).

Variety Is Big Data Integration Challenge #1
Often, the biggest Big Data integration challenge is variety, not volume. Consider all the different Big Data types that may require integration:

  • Massively Parallel Processing based Appliances - Examples include EMC Greenplum, HP Vertica, IBM Netezza, SAP Hana, and more
  • Columnar/tabular NoSQL Data Stores - Examples include Hadoop, Hypertable, and more
  • XML Document Data Stores - Examples include CouchDB, MarkLogic, and MongoDB, and more
  • Key/value Data Stores - Examples include Cassandra, Memcached, Voldemort, and more

Fortunately integrating heterogeneous data sources is the original raison d'etre of data virtualization.  Why do you think many still call it data federation?

Volume Is Big Data Integration Challenge #2
As listed above there are many ways to store and manage big data.  Similarly, a plethora of analysis tools exist such as MapR, Karmasphere, Alpine Data Labs and more.

The biggest volume challenge is how to query large data sets from these high-volume sources at speed in order to feed these analytics?

The answer is data virtualization.

Data virtualization platforms use sophisticated rule- and cost-based query-optimization strategies that automatically create a query plan that optimizes processing and performance, with minimum overhead.

Advanced Query Optimization Is the Key to Data Virtualization
Here are but a few of the query optimization strategies and techniques data virtualization provides:

  • Pushdown - Data virtualization offloads as much query processing as possible by pushing down select query operations such as string searches, comparisons, local joins, sorting, aggregating, grouping into the underlying data sources. Thus you can take advantage of native capabilities.
  • Parallel Processing - Data virtualization optimizes query execution by employing parallel and asynchronous request processing. After building an optimized query plan, the data virtualization server executes data service calls asynchronously on separate threads, reducing idle time and data source response latency.
  • Distributed Joins - Data virtualization detects when a query being executed involves data consumed from different data sources and tries to employ distributed query optimization techniques to improve overall performance and minimize the amount of data moved over the network.  A variety of sort-merge, semi, hash and nested-loop joins are leveraged depending on the nature of the query and data sources.
  • Caching - Data virtualization can be configured to cache results for query, procedure and web service calls.  When enabled, the caching engine stores the cached result sets and queries them as appropriate.
  • Advanced Query Optimization - Data virtualization provides a number of additional techniques and algorithms include data source grouping, join algorithm selection, join ordering, union-join inversion, predicate pooling and propagation, and projection pruning.
  • Integrated Network and Database Optimization - Even in a Big Data world; network bandwidth is generally the scarcest resource in the query processing pipeline. So reducing the amount of data that needs to be transferred has a significant impact on the latency and overall performance.  Data virtualization optimizes the network and the query processing capabilities of underlying big data sources intelligently, in combination.

Value and Velocity are Big Data Integration Challenges #3 and #4
Big Data itself only has value when the data is analyzed.  This analysis provides value by uncovering drivers for growth, finding better ways to attract and retain customers, and identifying opportunities for innovation and costs reduction.

As such the fastest path to Big Data analysis is also the fastest path to business value.

But everyone knows that providing analytics with the data required has always been difficult, with data integration long considered the biggest bottleneck in any analytics project.

The Data Warehousing Institute confirms this lack of agility.  Their recent study stated the average time needed to add a new data source to an existing BI application was 8.4 weeks in 2009, 7.4 weeks in 2010, and 7.8 weeks in 2011. And 33% of the organizations needed more than 3 months to add a new data source.

Data Virtualization Provides Velocity along with Analytic Value
According to Data Virtualization: Going Beyond Traditional Data Integration to Achieve Business Agility, data virtualization significantly accelerates data integration agility. Key to this success is data virtualization's

  • Streamlined data integration approach
  • Iterative development process
  • Adaptable change management process

Using data virtualization as a complement to existing data integration approaches, the ten organizations profiled in the book cut analytics project times in half or more.

This agility allowed the same teams to double their number of analytics projects, significantly accelerating the business value delivered.  In other words, value with velocity!

Variety, Volume, Velocity and Value
Big Data is all the rage.  And at first glance, the Big Data variety, volume, velocity and value challenges may seem extraordinarily difficult.

Proven technologies, such as data virtualization, provide proven approaches to addressing these "big" challenges.

So if Big Data is on your agenda, don't forget to make a big commitment to data virtualization.  You'll be glad you did.

More Stories By Robert Eve

Robert Eve is the EVP of Marketing at Composite Software, the data virtualization gold standard and co-author of Data Virtualization: Going Beyond Traditional Data Integration to Achieve Business Agility. Bob's experience includes executive level roles at leading enterprise software companies such as Mercury Interactive, PeopleSoft, and Oracle. Bob holds a Masters of Science from the Massachusetts Institute of Technology and a Bachelor of Science from the University of California at Berkeley.

@BigDataExpo Stories
SYS-CON Events announced today that Avere Systems, a leading provider of hybrid cloud enablement solutions, will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Avere Systems was created by file systems experts determined to reinvent storage by changing the way enterprises thought about and bought storage resources. With decades of experience behind the company’s founders, Avere got its ...
SYS-CON Events announced today that Taica will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. ANSeeN are the measurement electronics maker for X-ray and Gamma-ray and Neutron measurement equipment such as spectrometers, pulse shape analyzer, and CdTe-FPD. For more information, visit http://anseen.com/.
As you move to the cloud, your network should be efficient, secure, and easy to manage. An enterprise adopting a hybrid or public cloud needs systems and tools that provide: Agility: ability to deliver applications and services faster, even in complex hybrid environments Easier manageability: enable reliable connectivity with complete oversight as the data center network evolves Greater efficiency: eliminate wasted effort while reducing errors and optimize asset utilization Security: imple...
High-velocity engineering teams are applying not only continuous delivery processes, but also lessons in experimentation from established leaders like Amazon, Netflix, and Facebook. These companies have made experimentation a foundation for their release processes, allowing them to try out major feature releases and redesigns within smaller groups before making them broadly available. In his session at 21st Cloud Expo, Brian Lucas, Senior Staff Engineer at Optimizely, will discuss how by using...
In this strange new world where more and more power is drawn from business technology, companies are effectively straddling two paths on the road to innovation and transformation into digital enterprises. The first path is the heritage trail – with “legacy” technology forming the background. Here, extant technologies are transformed by core IT teams to provide more API-driven approaches. Legacy systems can restrict companies that are transitioning into digital enterprises. To truly become a lead...
SYS-CON Events announced today that Yuasa System will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Yuasa System is introducing a multi-purpose endurance testing system for flexible displays, OLED devices, flexible substrates, flat cables, and films in smartphones, wearables, automobiles, and healthcare.
With major technology companies and startups seriously embracing Cloud strategies, now is the perfect time to attend 21st Cloud Expo October 31 - November 2, 2017, at the Santa Clara Convention Center, CA, and June 12-14, 2018, at the Javits Center in New York City, NY, and learn what is going on, contribute to the discussions, and ensure that your enterprise is on the right path to Digital Transformation.
SYS-CON Events announced today that CAST Software will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. CAST was founded more than 25 years ago to make the invisible visible. Built around the idea that even the best analytics on the market still leave blind spots for technical teams looking to deliver better software and prevent outages, CAST provides the software intelligence that matter ...
SYS-CON Events announced today that Daiya Industry will exhibit at the Japanese Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Ruby Development Inc. builds new services in short period of time and provides a continuous support of those services based on Ruby on Rails. For more information, please visit https://github.com/RubyDevInc.
SYS-CON Events announced today that Evatronix will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Evatronix SA offers comprehensive solutions in the design and implementation of electronic systems, in CAD / CAM deployment, and also is a designer and manufacturer of advanced 3D scanners for professional applications.
As businesses evolve, they need technology that is simple to help them succeed today and flexible enough to help them build for tomorrow. Chrome is fit for the workplace of the future — providing a secure, consistent user experience across a range of devices that can be used anywhere. In her session at 21st Cloud Expo, Vidya Nagarajan, a Senior Product Manager at Google, will take a look at various options as to how ChromeOS can be leveraged to interact with people on the devices, and formats th...
First generation hyperconverged solutions have taken the data center by storm, rapidly proliferating in pockets everywhere to provide further consolidation of floor space and workloads. These first generation solutions are not without challenges, however. In his session at 21st Cloud Expo, Wes Talbert, a Principal Architect and results-driven enterprise sales leader at NetApp, will discuss how the HCI solution of tomorrow will integrate with the public cloud to deliver a quality hybrid cloud e...
SYS-CON Events announced today that Taica will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Taica manufacturers Alpha-GEL brand silicone components and materials, which maintain outstanding performance over a wide temperature range -40C to +200C. For more information, visit http://www.taica.co.jp/english/.
When it comes to cloud computing, the ability to turn massive amounts of compute cores on and off on demand sounds attractive to IT staff, who need to manage peaks and valleys in user activity. With cloud bursting, the majority of the data can stay on premises while tapping into compute from public cloud providers, reducing risk and minimizing need to move large files. In his session at 18th Cloud Expo, Scott Jeschonek, Director of Product Management at Avere Systems, discussed the IT and busine...
Organizations do not need a Big Data strategy; they need a business strategy that incorporates Big Data. Most organizations lack a road map for using Big Data to optimize key business processes, deliver a differentiated customer experience, or uncover new business opportunities. They do not understand what’s possible with respect to integrating Big Data into the business model.
Enterprises have taken advantage of IoT to achieve important revenue and cost advantages. What is less apparent is how incumbent enterprises operating at scale have, following success with IoT, built analytic, operations management and software development capabilities – ranging from autonomous vehicles to manageable robotics installations. They have embraced these capabilities as if they were Silicon Valley startups. As a result, many firms employ new business models that place enormous impor...
Amazon is pursuing new markets and disrupting industries at an incredible pace. Almost every industry seems to be in its crosshairs. Companies and industries that once thought they were safe are now worried about being “Amazoned.”. The new watch word should be “Be afraid. Be very afraid.” In his session 21st Cloud Expo, Chris Kocher, a co-founder of Grey Heron, will address questions such as: What new areas is Amazon disrupting? How are they doing this? Where are they likely to go? What are th...
SYS-CON Events announced today that MIRAI Inc. will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. MIRAI Inc. are IT consultants from the public sector whose mission is to solve social issues by technology and innovation and to create a meaningful future for people.
SYS-CON Events announced today that Dasher Technologies will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Dasher Technologies, Inc. ® is a premier IT solution provider that delivers expert technical resources along with trusted account executives to architect and deliver complete IT solutions and services to help our clients execute their goals, plans and objectives. Since 1999, we'v...
Though cloud is the future of enterprise computing, a smooth transition of legacy applications and systems is critical for seamless business operations. IT professionals are eager to start leveraging the cost, scale and other benefits of cloud, but with massive investments already in place in existing infrastructure and a number of compliance and resource hurdles, it can be challenging to move to a cloud-based infrastructure.