Welcome!

@DXWorldExpo Authors: Pat Romanski, Yeshim Deniz, SmartBear Blog, Elizabeth White, Liz McMillan

Related Topics: @DXWorldExpo, Java IoT, Open Source Cloud, Containers Expo Blog, Agile Computing, @CloudExpo

@DXWorldExpo: Article

Big Data: So What! That’s Why You Virtualize

How data virtualization enables Big Data volume, variety, velocity and value

Big Data!  Yes it's BIG!

The volume is BIG!  The variety is BIG!  The velocity is BIG!

And hopefully the business value is BIG!

New Opportunities Bring New Ways to Leverage Proven Technology
There is no shortage of media articles, analyst reports, tradeshows, blogs and other source of Big Data technology insight and advice.

But it strikes me that in our search to be on the leading edge, we may be overlooking some great existing technology.

In fact, some technology, for example data virtualization, is even more useful in a Big Data world.

What Is Data Virtualization?
Data virtualization is an agile data integration approach organizations use to gain more insight from their data.  This includes traditional sources such as transaction systems, data warehouses and more as well as new sources such the cloud and Big Data.

Unlike data consolidation or data replication, data virtualization integrates these diverse data types without costly extra copies and additional data management complexity.  Seriously, if the data is already big, why make it even bigger by copying and storing it again and again?

With data virtualization, you respond faster to ever changing analytics and BI needs, fast-track your data management evolution and save 50-75% over data replication and consolidation.   In other words, you deliver value, the most important V but often not listed with the 3 Vs of Big Data (Volume, Velocity & Variety).

Variety Is Big Data Integration Challenge #1
Often, the biggest Big Data integration challenge is variety, not volume. Consider all the different Big Data types that may require integration:

  • Massively Parallel Processing based Appliances - Examples include EMC Greenplum, HP Vertica, IBM Netezza, SAP Hana, and more
  • Columnar/tabular NoSQL Data Stores - Examples include Hadoop, Hypertable, and more
  • XML Document Data Stores - Examples include CouchDB, MarkLogic, and MongoDB, and more
  • Key/value Data Stores - Examples include Cassandra, Memcached, Voldemort, and more

Fortunately integrating heterogeneous data sources is the original raison d'etre of data virtualization.  Why do you think many still call it data federation?

Volume Is Big Data Integration Challenge #2
As listed above there are many ways to store and manage big data.  Similarly, a plethora of analysis tools exist such as MapR, Karmasphere, Alpine Data Labs and more.

The biggest volume challenge is how to query large data sets from these high-volume sources at speed in order to feed these analytics?

The answer is data virtualization.

Data virtualization platforms use sophisticated rule- and cost-based query-optimization strategies that automatically create a query plan that optimizes processing and performance, with minimum overhead.

Advanced Query Optimization Is the Key to Data Virtualization
Here are but a few of the query optimization strategies and techniques data virtualization provides:

  • Pushdown - Data virtualization offloads as much query processing as possible by pushing down select query operations such as string searches, comparisons, local joins, sorting, aggregating, grouping into the underlying data sources. Thus you can take advantage of native capabilities.
  • Parallel Processing - Data virtualization optimizes query execution by employing parallel and asynchronous request processing. After building an optimized query plan, the data virtualization server executes data service calls asynchronously on separate threads, reducing idle time and data source response latency.
  • Distributed Joins - Data virtualization detects when a query being executed involves data consumed from different data sources and tries to employ distributed query optimization techniques to improve overall performance and minimize the amount of data moved over the network.  A variety of sort-merge, semi, hash and nested-loop joins are leveraged depending on the nature of the query and data sources.
  • Caching - Data virtualization can be configured to cache results for query, procedure and web service calls.  When enabled, the caching engine stores the cached result sets and queries them as appropriate.
  • Advanced Query Optimization - Data virtualization provides a number of additional techniques and algorithms include data source grouping, join algorithm selection, join ordering, union-join inversion, predicate pooling and propagation, and projection pruning.
  • Integrated Network and Database Optimization - Even in a Big Data world; network bandwidth is generally the scarcest resource in the query processing pipeline. So reducing the amount of data that needs to be transferred has a significant impact on the latency and overall performance.  Data virtualization optimizes the network and the query processing capabilities of underlying big data sources intelligently, in combination.

Value and Velocity are Big Data Integration Challenges #3 and #4
Big Data itself only has value when the data is analyzed.  This analysis provides value by uncovering drivers for growth, finding better ways to attract and retain customers, and identifying opportunities for innovation and costs reduction.

As such the fastest path to Big Data analysis is also the fastest path to business value.

But everyone knows that providing analytics with the data required has always been difficult, with data integration long considered the biggest bottleneck in any analytics project.

The Data Warehousing Institute confirms this lack of agility.  Their recent study stated the average time needed to add a new data source to an existing BI application was 8.4 weeks in 2009, 7.4 weeks in 2010, and 7.8 weeks in 2011. And 33% of the organizations needed more than 3 months to add a new data source.

Data Virtualization Provides Velocity along with Analytic Value
According to Data Virtualization: Going Beyond Traditional Data Integration to Achieve Business Agility, data virtualization significantly accelerates data integration agility. Key to this success is data virtualization's

  • Streamlined data integration approach
  • Iterative development process
  • Adaptable change management process

Using data virtualization as a complement to existing data integration approaches, the ten organizations profiled in the book cut analytics project times in half or more.

This agility allowed the same teams to double their number of analytics projects, significantly accelerating the business value delivered.  In other words, value with velocity!

Variety, Volume, Velocity and Value
Big Data is all the rage.  And at first glance, the Big Data variety, volume, velocity and value challenges may seem extraordinarily difficult.

Proven technologies, such as data virtualization, provide proven approaches to addressing these "big" challenges.

So if Big Data is on your agenda, don't forget to make a big commitment to data virtualization.  You'll be glad you did.

More Stories By Robert Eve

Robert Eve is the EVP of Marketing at Composite Software, the data virtualization gold standard and co-author of Data Virtualization: Going Beyond Traditional Data Integration to Achieve Business Agility. Bob's experience includes executive level roles at leading enterprise software companies such as Mercury Interactive, PeopleSoft, and Oracle. Bob holds a Masters of Science from the Massachusetts Institute of Technology and a Bachelor of Science from the University of California at Berkeley.

@BigDataExpo Stories
"This week we're really focusing on scalability, asset preservation and how do you back up to the cloud and in the cloud with object storage, which is really a new way of attacking dealing with your file, your blocked data, where you put it and how you access it," stated Jeff Greenwald, Senior Director of Market Development at HGST, in this SYS-CON.tv interview at 18th Cloud Expo, held June 7-9, 2016, at the Javits Center in New York City, NY.
In a world where the internet rules all, where 94% of business buyers conduct online research, and where e-commerce sales are poised to fall between $427 billion and $443 billion by the end of this year, we think it's safe to say that your website is a vital part of your business strategy. Whether you're a B2B company, a local business, or an e-commerce site, digital presence is key to maintain in your drive towards success. Digital Performance will take priority in 2018 for the following reason...
Rodrigo Coutinho is part of OutSystems' founders' team and currently the Head of Product Design. He provides a cross-functional role where he supports Product Management in defining the positioning and direction of the Agile Platform, while at the same time promoting model-based development and new techniques to deliver applications in the cloud.
Business professionals no longer wonder if they'll migrate to the cloud; it's now a matter of when. The cloud environment has proved to be a major force in transitioning to an agile business model that enables quick decisions and fast implementation that solidify customer relationships. And when the cloud is combined with the power of cognitive computing, it drives innovation and transformation that achieves astounding competitive advantage.
In his session at Cloud Expo, Alan Winters, U.S. Head of Business Development at MobiDev, presented a success story of an entrepreneur who has both suffered through and benefited from offshore development across multiple businesses: The smart choice, or how to select the right offshore development partner Warning signs, or how to minimize chances of making the wrong choice Collaboration, or how to establish the most effective work processes Budget control, or how to maximize project result...
"Software-defined storage is a big problem in this industry because so many people have different definitions as they see fit to use it," stated Peter McCallum, VP of Datacenter Solutions at FalconStor Software, in this SYS-CON.tv interview at 18th Cloud Expo, held June 7-9, 2016, at the Javits Center in New York City, NY.
In his keynote at 19th Cloud Expo, Sheng Liang, co-founder and CEO of Rancher Labs, discussed the technological advances and new business opportunities created by the rapid adoption of containers. With the success of Amazon Web Services (AWS) and various open source technologies used to build private clouds, cloud computing has become an essential component of IT strategy. However, users continue to face challenges in implementing clouds, as older technologies evolve and newer ones like Docker c...
Data is the fuel that drives the machine learning algorithmic engines and ultimately provides the business value. In his session at Cloud Expo, Ed Featherston, a director and senior enterprise architect at Collaborative Consulting, discussed the key considerations around quality, volume, timeliness, and pedigree that must be dealt with in order to properly fuel that engine.
As organizations shift towards IT-as-a-service models, the need for managing and protecting data residing across physical, virtual, and now cloud environments grows with it. Commvault can ensure protection, access and E-Discovery of your data – whether in a private cloud, a Service Provider delivered public cloud, or a hybrid cloud environment – across the heterogeneous enterprise. In his general session at 18th Cloud Expo, Randy De Meno, Chief Technologist - Windows Products and Microsoft Part...
Andi Mann, Chief Technology Advocate at Splunk, is an accomplished digital business executive with extensive global expertise as a strategist, technologist, innovator, marketer, and communicator. For over 30 years across five continents, he has built success with Fortune 500 corporations, vendors, governments, and as a leading research analyst and consultant.
"Cloud computing is certainly changing how people consume storage, how they use it, and what they use it for. It's also making people rethink how they architect their environment," stated Brad Winett, Senior Technologist for DDN Storage, in this SYS-CON.tv interview at 20th Cloud Expo, held June 6-8, 2017, at the Javits Center in New York City, NY.
In his session at 20th Cloud Expo, Brad Winett, Senior Technologist for DDN Storage, will present several current, end-user environments that are using object storage at scale for cloud deployments including private cloud and cloud providers. Details on the top considerations of features and functions for selecting object storage will be included. Brad will also touch on recent developments in tiering technologies that deliver single solution and an end-user view of data across files and objects...
In his keynote at 18th Cloud Expo, Andrew Keys, Co-Founder of ConsenSys Enterprise, provided an overview of the evolution of the Internet and the Database and the future of their combination – the Blockchain. Andrew Keys is Co-Founder of ConsenSys Enterprise. He comes to ConsenSys Enterprise with capital markets, technology and entrepreneurial experience. Previously, he worked for UBS investment bank in equities analysis. Later, he was responsible for the creation and distribution of life settl...
No hype cycles or predictions of zillions of things here. IoT is big. You get it. You know your business and have great ideas for a business transformation strategy. What comes next? Time to make it happen. In his session at @ThingsExpo, Jay Mason, Associate Partner at M&S Consulting, presented a step-by-step plan to develop your technology implementation strategy. He discussed the evaluation of communication standards and IoT messaging protocols, data analytics considerations, edge-to-cloud tec...
In his session at @ThingsExpo, Dr. Robert Cohen, an economist and senior fellow at the Economic Strategy Institute, presented the findings of a series of six detailed case studies of how large corporations are implementing IoT. The session explored how IoT has improved their economic performance, had major impacts on business models and resulted in impressive ROIs. The companies covered span manufacturing and services firms. He also explored servicification, how manufacturing firms shift from se...
IoT is at the core or many Digital Transformation initiatives with the goal of re-inventing a company's business model. We all agree that collecting relevant IoT data will result in massive amounts of data needing to be stored. However, with the rapid development of IoT devices and ongoing business model transformation, we are not able to predict the volume and growth of IoT data. And with the lack of IoT history, traditional methods of IT and infrastructure planning based on the past do not app...
Organizations planning enterprise data center consolidation and modernization projects are faced with a challenging, costly reality. Requirements to deploy modern, cloud-native applications simultaneously with traditional client/server applications are almost impossible to achieve with hardware-centric enterprise infrastructure. Compute and network infrastructure are fast moving down a software-defined path, but storage has been a laggard. Until now.
Digital Transformation is much more than a buzzword. The radical shift to digital mechanisms for almost every process is evident across all industries and verticals. This is often especially true in financial services, where the legacy environment is many times unable to keep up with the rapidly shifting demands of the consumer. The constant pressure to provide complete, omnichannel delivery of customer-facing solutions to meet both regulatory and customer demands is putting enormous pressure on...
The best way to leverage your CloudEXPO | DXWorldEXPO presence as a sponsor and exhibitor is to plan your news announcements around our events. The press covering CloudEXPO | DXWorldEXPO will have access to these releases and will amplify your news announcements. More than two dozen Cloud companies either set deals at our shows or have announced their mergers and acquisitions at CloudEXPO. Product announcements during our show provide your company with the most reach through our targeted audienc...
With 10 simultaneous tracks, keynotes, general sessions and targeted breakout classes, @CloudEXPO and DXWorldEXPO are two of the most important technology events of the year. Since its launch over eight years ago, @CloudEXPO and DXWorldEXPO have presented a rock star faculty as well as showcased hundreds of sponsors and exhibitors!