Showing posts with label data quality. Show all posts
Showing posts with label data quality. Show all posts

Tuesday, March 16, 2021

ActiveNav Automates Data Inventory Updates

Our on-going tour of privacy systems has already included stops at BigID  and Trust-Hub, which both build inventories of customer data. Apparently I’m drawn the topic, since I recently found myself looking at ActiveNav, which turns out to be yet another data inventory system. It’s different enough from the others to be worth a review of its own. So here goes.

Like other inventory systems, ActiveNav builds a map of data stored in company systems. In ActiveNav’s case, this can literally be a geographic map showing the location of data centers, starting from a global view and drilling down to regions, cities, and sites. Users can also select other views, including business units and repository types. The lowest level in each view is a single container, whose contents ActiveNav will automatically explore by reading metadata attributes such as field names and data formats.

The system applies rules and keywords to the metadata to determine the type of data stored in each field, without reading the actual file contents. (A supplementary module that allows content examination is due for release soon.)  It stores its findings in its own repository, again without copying any actual information – so there’s no worry about data breaches from ActiveNav itself.

One disadvantage of ActiveNav’s approach is that relying only on metadata limits the chances of finding sensitive information that is not labeled accurately, something that BigID does especially well. Similarly, ActiveNav doesn’t map relations between data stored in different containers, so it cannot build a company-wide data model. This is a strength of Trust-Hub.

Still, ActiveNav’s ability to explore and classify data repositories without human guidance is a major improvement over manually-built data inventories. Its second big benefit is a “data health” score based on its findings. This is calculated for each container with scores for factors including: risk, including intellectual property and security issues; privacy compliance, based on presence of IDs and other data types; and data quality, including duplicate, obsolete, stale and trivial contents. Scores for each container are combined to create scores for repositories, locations, business units, and other higher levels. This gives users a quick way to find problem areas and track data health over time.

ActiveNav addresses what may be the biggest data inventory pain of all: keeping information up-to-date. The system automates the update process by receiving continuous notifications of metadata changes from systems that are set up to send them. In other cases, ActiveNav can query repositories to look for metadata that has been updated since its last visit. Of course, this requires providing the system with credentials to access that information.

ActivNav was founded in 2008. Until recently, it offered only a conventional on-premise software license with one-time costs starting around $100,000. This is sold this primarily through partners who work on data management projects for heavily regulated industries and governments. The company has recently introduced a SaaS version of its data inventory system that starts at $10,000 per year. It also offers data governance and compliance modules.

Monday, August 22, 2016

ABM Vendor Guide: What to Look for in External Data Sources

Last week’s posts introduced our new Raab Guide to ABM Vendors (buy it here) and introduced a framework four process ABM steps, six system functions, and six key sub-functions. The idea was that functions define major categories of systems, while the sub-functions differentiate systems within each category. The world isn’t really quite this simple, if only because many systems provide more than one function. But the sub-functions are still important for stack design and vendor selection.

My plan this week is to follow up with a sequence of posts that go through each sub-function in some depth.  Let’s start with the first sub-function, External Data. 

ABM Process
System Function
Sub-Function
Number of Vendors
Identify Target Accounts
Assemble Data
External Data
28
Select Targets
Target Scoring
15
Plan Interactions
Assemble Messages
Customized Messages
6
Select Messages
State-Based Flows
10
Execute Interactions
Deliver Messages
Execution
19
Analyze Results
Reporting
Result Analysis
16


Vendors that support this sub-function gather account and contact information from the Internet, private, and government sources and purchase it from other vendors. They may resell the data to marketers or use it themselves to support tasks such as account scoring or ad targeting. 

(To put things in a broader context, “external data” can be contrasted with “internal data”, which comes from a company’s own systems for CRM, marketing automation, Web analytics, order processing, customer support, etc. Internal data is most important later in the sales cycle, when prospects and customers are interacting with the company directly. External data is most important at the start, when the company hasn’t identified its target accounts or established direct relationships with them.)

External data may seem like a commodity – after all, all vendors have access to pretty much the same sources. Yet there’s probably more variety among the vendors in this category than any other. Some key differentiators identified in the ABM Guide include:

  • types of data provided (companies, contacts, events, intent, technology used)
  • number and types of data sources (company Web pages, publisher Web pages, ad exchanges and networks, job posting sites, social networks, IP directories, financial reports, government files, industry and professional directories, news feeds, etc.  Different sources provide different data types.)
  • depth of data (to measure this, get a list of data elements)
  • quality of data (harder to measure: review some sample records and have the vendor explain their quality methods)
  • coverage by region, company size, industry, etc. (depends heavily on data types and sources)
  • coverage by language (many systems extract data using natural language technology that can only read English)
  • how often data is refreshed (which involves two issues: how often are sources revisited and how quickly do changes get communicated to clients)
  • on-demand updates for individual accounts or contacts (to get up-to-the minute information on a new or existing account)
  • add new data sources to meet specific client needs (e.g., reports of new research contracts in the client's industry)
  • custom research to supplement public information (in particular, some vendors do custom research to identify the IP addresses used by target accounts)
  • custom taxonomies for intent analysis (because standard taxonomies may not be precise enough for specialized client needs)
  • maturity of data management processes (how long they’ve been evolving, size of team, etc.)
  • data verification methods (phone call, test for email bounces, compare against other sources, etc.)
  • special methods to associate personal and business emails, attach leads to accounts, find social media handles, etc. (vendors may do different kinds of “fuzzy” matching, machine learning, or natural language processing to uncover or infer relationships when exact matches are not available)
  • load client data and match against it for enhancement (most vendors will do this but some require the client to do its own matching)
  • continuous updates of client data (reporting on changes as the vendor learns about them; requires uploading a list of accounts or individuals to monitor)
  • provide personal identifiers on contact records (name, address, phone, email address, social media handle, etc.; not all vendors do this, especially in countries with strict privacy laws; different identifiers are also treated differently)
  • provide a complete universe of all companies in a target market (some vendors only enhance records already in the client database, others provide "net new" records as well.)
  • find social connections between company employees and target account employees (make sure this is done without violating the social network terms of service).
  • real-time processes to identify Web site visitors, auto-fill Web forms or verify form entries, show data to sales people, support ad targeting, etc.
  • add ownership relationships to accounts (headquarters/branch, parent/subsidiary, brand/franchisee, etc.) in general and to the D-U-N-S Number in particular
  • fee structure (most are vendors charge per record and/or based on the data types; some are performance-based)

Which of these are important will depend on your business needs and approach to ABM. For example, if you sell to small businesses, then coverage is critical because many vendors identify companies using IP address or Web domain – things many small businesses do not have. On the other hand, if you want to target Web messages to large enterprises, your critical need will be real-time identification of Web site visitors, something only a few vendors can support.

The issue hovering over all this is data quality. If quality is poor, then nothing else matters. Quality can be a somewhat tricky concept, since it’s not just accuracy or coverage or currency.  My personal favorite definition of quality is “fitness for purpose”, which makes the point that the quality of data (or anything else) is related to how you’re going to use it. But even assuming you know exactly what you need, you can’t predict quality based on a checklist of features or attributes. The only practical approach is to get some sample data and see how it performs, whether by comparing it to known correct data, testing it directly via phone calls or surveys, or using it in a marketing program and measuring the results. Experienced data-driven marketers have known this forever, but less experienced marketers may not realize that all data isn’t as good as they’d like to assume. There’s not much I can do in the ABM Guide to solve this issue, but smart marketers can use the Guide to identify data vendors who meet their other requirements, and then test those vendors' data to ensure the quality is what they need.




Monday, January 25, 2016

Avention DataVision Gives Sales and Marketing Systems Unified Access to B2B Customer Data Quality and Alerts

My look last week at True Influence’s InsightBASE, a relatively new-fangled approach to intent data, was karmically balanced by a conversation with Avention, a old-line data aggregator that traces its roots to CD-ROM business lists from Lotus OneSource. The folks at Avention had reached out to discuss their latest product, DataVision, which extends Avention’s reach from sales enablement to marketing systems.  The goal is giving clients a single data source to support both departments.

DataVision lets clients upload customer lists to be cleaned and enhanced by matching against Avention’s own master file, which is itself compiled from some seventy sources. Sales and marketing systems can then access the results in an online database, providing all departments with a single, consistent view of their consolidated data. The information includes both companies and contacts and supplements standard profile information with event-based "signals" derived from news reports, company Web sites, and social media postings. Clients can set up alerts based on signals and can acquire new names that are similar to their current customers.

If this sounds familiar, it’s because Reachforce, InsideView, SalesLoft and other data vendors offer similar services. Predictive modeling vendors including Leadspace, Lattice Engines, Mintigo, and Everstring also provide enhancement and signal-based alerts, although usually with less depth of detail. The biggest difference is those vendors usually send the enhanced information back to client systems rather than keeping it in an external database which sales and marketing systems access directly.

But different isn’t necessarily better. No one will discard their CRM or marketing automation database and use the DataVision file instead. There’s simply too much other information within the sales and marketing systems. So, in practice, DataVision will be used to update a company’s existing databases, pretty much the same as its competitors. The data may be a bit fresher, since any query to DataVision will return the latest information available to Avention. DataVision also provides some nice tools to visualize the distribution of a client’s customers across geography, industry, company size, and other dimensions, and to compare those distributions with the entire Avention universe of known firms. Again, these features are useful even if they are not necessarily unique.

In short, Avention DataVision is a solid option when you’re looking to clean and enhance your company’s customer and prospect data – something every firm needs to do. Intent data and predictive modeling are not part of the mix yet, but it’s easy to imagine those being added in the future. Whether Avention is your best choice will depend on your specific situation.  The only way to know is to define your exact requirements, test several sources, and evaluate the results. The good news is you have lots of vendors to choose from, so you have a good chance of finding one that fits your needs.

Tuesday, July 07, 2015

Does Future Marketing Technology Require Perfect Data?


I mentioned in my last post that I’ve started to think in terms of three realities: today (the next two years), tomorrow (two to five years out), and later (after five years). Like the famous New Yorker magazine cover that showed a detailed knowledge of Manhattan and increasingly vague view of more distant regions, our picture of the immediate future is much more nuanced than what happens farther out. One result is an apparent assumption that future technology will work much better than today’s technology – not because anyone really thinks that future technology will be perfect, but because we can’t see where its imperfections will appear.

I’ve been thinking about this because so my own predictions are premised on increasingly detailed knowledge about customers and prospects. Both the “madtech” vision of broad access to third-party data and the “robotech” vision of delegating decisions to machines assume that effectively complete data will be available about each customer. But a quick look at today’s data shows that is far from true. Here are some factoids I’ve been gathering to illustrate the point:


- 37% of mobile ad locations are accurate to within 100 meters (Thinknear)

- 30-55% match rates for B2C individual-level onboarding (LiveRamp)

- 16-29% match rates for B2B individual-level data enrichment: (Raab Associates client tests)

- 14% match rates and low predictive value for B2B account-level intent data: (Infer)

And this doesn’t even begin to address predictive modeling, where even a 10x lift vs average still implies many errors at the individual level.

Contemplating these results does give me pause. At some point, poor data means that theoretically possible approaches are not practical because of low coverage or insufficient performance. Those constraints won’t magically vanish in the future, even though they’re not visible at this distance.

Being a technology optimist, I assume that data will get better over time. But I can’t cite much evidence to support my optimism.  If anything, the number of new data sources is outstripping improvements in existing sources. The true core challenge is identity resolution, which means associating data from different sources with the right individual profile. Cross-device matching is the current focus of this discussion but covers just part of the problem.

It’s a safe bet that perfect data won’t be available in two years or five years or probably ever. But the real question is whether enough good data will be available to support the futures I’ve been forecasting.

I think a realistic view is that some data will be more available than other data, and, as a result, some portions of the visions will happen while others do not. Customer data is likely to be richer than prospect data, since customers will grant permission to link with external data sources (or take actions that make linking easier even without their permission). Sharing among complementary companies – for example, airlines and hotels – will be easier to negotiate than sharing with anyone through public exchanges. Data about objects, such as cars or groceries or homes, should be less sensitive than data about individuals (even though there’s obviously a close relationship between objects and their owners). Data about public behaviors, such as travel and store visits, is less sensitive than data about private matters such as health care.  (See this recent Altimeter Group report for more information on consumer attitudes to privacy.)

In short, the future will remain unevenly distributed, as William Gibson observed. Marketers and the technologists who support them need both the ideal vision of how things would work in a world of perfect data (which isn’t the same as a perfect world!) and the realistic understanding of what’s likely to be practical within their planning horizon. They can then aggressively pursue opportunities revealed by the vision without chasing chimeras that will never appear. This pursuit is essential: tomorrow always comes, but the future won’t happen by itself.




Monday, December 01, 2014

Radius Provides High Quality Data on Small Businesses

When I first spoke with Radius just over one year ago, the company had already pivoted from its initial concept as a mobile app to connect consumers with local business events, to building a comprehensive list of small businesses and their attributes. Fast-forward twelve months and the company has again adjusted its offering, now presenting itself as a “marketing intelligence platform” that helps business marketers find prospects who are similar to their current buyers. This latest vision was appealing enough to attract $54.7 million in funding in September, bringing the announced total to over $80 million. So I’m guessing Radius will stick with this approach for a while.


What Radius does will sound broadly familiar to loyal readers of this blog: it scans social media, Web pages, government records, and other online sources to build a list of more than 20 million U.S. businesses and their attributes. It supplements these with conventional data sources to capture businesses with a limited digital profile. In its current incarnation, Radius also imports a list of won and lost deals from each client’s CRM system (direct connection to Salesforce.com, batch imports from others) and shows how well each attribute correlates with success.

Users can review the attribute list, create segments based on attributes, and analyze the attributes of each segment as a group. They can also flag existing segment members within the client’s current CRM database and import segment members who are not already in the client’s CRM (a.k.a. “net new prospects”). The imported records include basic company information and other attributes the client has preselected, but the system will not correct or enhance existing CRM records. Segment membership is adjusted automatically as Radius updates its data, which happens weekly. The system does not store fixed lists of segment members at a point in time, although users could achieve this by tagging records in CRM as they are imported. 

And that’s pretty much it. No list of the most important attributes, no predictive modeling, one contact name per company, no alerts based on buying signals, no campaign analysis: just company information compared to your own customers, an way to build segments, and an option to purchase new prospects. The company plans to address some of these gaps but has not released the details.

Whatever its limits, Radius has attracted some big-name customers, most notably American Express, as well as all that funding. The primary reason seems to be data quality: the company says it can usually match 80% to 90% of the businesses in a well-maintained CRM system and that client tests have shown it is more accurate than competitors. This is both impressive and important, especially where small businesses are concerned. Available data includes basics (address, phone, industry, company size, revenue, contact name), Web activity (presence of a Web site, Facebook and Twitter accounts, use of daily deals and check ins, and average review ratings), and technologies used.

The system has some other advantages.  New clients are deployed in 24 hours, including the time to import CRM data and calculate success rates by attribute. The user interface is attractive and intuitive.  Pricing starts at $15,000 per year for small enterprises.  It is based on the number of company records in the client database, so it doesn’t increase based on how heavily the system is used.  It also includes use of the system software and credits for some number of new prospects imported from the Radius database.

In short, Radius strikes me as a solid solution for what it does, which is provide targeted company-level prospect lists and profiles of your current customer base. If that’s what you want, take a closer look. If you want to know more about trigger events or individual contacts or want lead scoring or other types of predictive modeling, you’ll probably be happier with something else.

Friday, December 13, 2013

Webinar, December 18: How Marketers Can (Finally) Get Good Customer Data

Let’s face it: no real work will get done next week, what with all the holiday parties and caroling and so forth. So you might as well set aside 2:00 to 3:00 p.m. Eastern time on Wednesday, December 18 and register for the Webinar I’m co-presenting with RedPoint Global on Customer Data Platforms.

In addition to uncovering the secret relationship between cuneiform and Justin Bieber,


you’ll learn about our latest discovery: a new species of Customer Data Platform, bringing the known total to four. We’ll even provide a handy field guide to identifying which is which. Join us, and gain enough new information to fuel your party conversations for the rest of the holiday season!

Thursday, September 19, 2013

New Study: Three Types of Customer Data Platform Address Cross-Channel Marketing Needs

My detailed study of Customer Data Platforms should be released next week. Now that the information is assembled, I can at last pull back and get a good overview of what I’ve found.

Perhaps the most interesting discovery has been that the CDP vendors cluster into three main groups.

• B2B data enhancement. These build a large reference database of companies and employees, which they match against records imported from their clients. They generally return corrected and enhanced data and lead scores based on models built from the client’s customer files. Their reference databases are built from multiple public, commercial, and proprietary sources, and are assembled using sophisticated matching engines. Most also perform their own scans of Web sites and social networks to extract sales-relevant information such as technology use and changes that suggest buying opportunities. These vendors vary considerably in the data they return, ranging from lead scores only to recommended marketing treatments to full customer profiles. Some also provide prospect lists of companies that are not already in the client’s own database. CDP vendors in this group include Infer, Lattice Engines, Mintigo, and ReachForce.

These systems compete with non-CDP products which also add or enhance prospect records but do not maintain a database with their clients’ customers. These include Web scanning systems such as InsideView, LeadSpace, and SalesLoft, and general data compilers including NetProspex, Demandbase, Data.com, ZoomInfo, and OneSource. The predictive modeling features also compete to some degree with end-user-oriented marketing analytics and modeling software such as Birst, GoodData, Cloud9 Analytics, AutoBox, and Predixion Software. Data cleansing competitors include services from firms such as D&B, as well as data management software for technical users such as Informatica, Experian QAS, and FullContact.

• Campaigns. These systems build a multi-source marketing database from the client’s own data and either recommend marketing treatments to execution systems or execute marketing campaigns directly. These are primarily used for consumer marketing although they also have B2B clients. Most have sophisticated matching capabilities. This group includes Silverpop with its Universal Behavior feature, NICE’s Causata, AgilOne, and RedPoint.

This group competes with conventional consumer marketing automation products, which provide similar campaign management abilities but lack the CDPs' database flexibility, database management, and customer matching features.

• Audience management. These systems build a database of customers and their responses to online display advertisements. They then build models that predict the customers’ probability of responding to future advertisements and provide recommendations for how much to bid and which content to display. These systems perform the same basic functions as standard online audience management systems (Data Management Platforms, or DMPs) and provide the same very quick responses needed for real time bidding (usually under 100 milliseconds). The major difference is that they also recommend messages in other channels, such as Web site personalization or email campaigns. Like DMPs, they work primarily at the Web cookie level, can link cookies known to relate to the same customer, and can be linked to actual customer names and addresses in external systems. This group includes IgnitionOne, [x+1], and Knotice.

This group overlaps with recommendation and ad targeting engines and DMP systems. Those products provide similar functions but do not track identified individuals and are often limited to single channel executions.

Given that each group addresses a different business need, you might wonder why I think they should all be lumped together under the CDP label. Quite simply, it’s because they are all addressing a portion of the same larger problem, which is how marketers can get a complete view of their customers and use that view to coordinate treatments across channels. What marketers truly need is a combination of the features from each group: data enhancement from external sources, for consumers as well as B2B; sophisticated customer matching and treatment selection; and integration of online advertising audiences with traditional customer databases. Each of these systems has the potential to grow into a complete solution, and the normal dynamics of software industry growth will push them towards pursuing that potential. So I expect the categories to overlap increasingly over the next few years and eventually merge into complete Customer Data Platforms as I envision them.

Incidentally and tangentially related: I'll be giving a Webinar with ReachForce on October 2 on Data Quality for Hipsters, a name that started as a joke but does make the point that data quality is essential for cutting-edge marketing.  YOLO, so you might as well attend.  I'm already working on the mustache.



Thursday, February 14, 2008

What's New at DataFlux? I Thought You'd Never Ask.

What with it being Valentine’s Day and all, you probably didn’t wake up this morning asking yourself, “I wonder what’s new with DataFlux?” That, my friend, is where you and I differ. Except that I actually asked myself that question a couple of weeks ago, and by now have had time to get an answer. Which turns out to be rather interesting.

DataFlux, as anyone still reading this probably knew already, is a developer of data quality software and is owned by SAS. DataFlux’s original core technology was a statistical matching engine that automatically analyzes input files and generates sophisticated keys which are similar for similar records. This has now been supplemented by a variety of capabilities for data profiling, analysis, standardization and verification, using reference data and rules in addition to statistical methods. The original matching engine is now just one component within a much larger set of solutions.

In fact, and this is what I find interesting, much of DataFlux’s focus is now on the larger issue of data governance. This has more to do with monitoring data quality than simple matching. DataFlux tells me the change has been driven by organizations that face increasing pressures to prove they are doing a good job with managing their data, for reasons such as financial reporting and compliance with government regulations.

The new developments also encompass product information and other types of non-name and address data, usually labeled as “master data management”. DataFlux reports that non-customer data is is the fastest growing portion of its business. DataFlux is well suited for non-traditional matching applications because the statistical approach does not rely on topic-specific rules and reference bases. Of course, DataFlux does use rules and reference information when appropriate.

The other recent development at DataFlux has been creation of “accelerators”, which are prepackaged rules, processes and reports for specific tasks. DataFlux started offering these in 2007 and now lists one each for customer data quality, product data quality, and watchlist compliance. More are apparently on the way. Applications like this are a very common development in a maturing industry, as companies that started by providing tools gain enough experience to understand how the applications commonly built with those tools. The next step—which DataFlux hasn’t reached yet—is to become even more specific by developing packages for particular industries. The benefit of these applications is that they save clients work and allow quicker deployments.

Back to governance. DataFlux’s movement in that direction is an interesting strategy because it offers a possible escape from the commoditization of its core data quality functions. Major data quality vendors including Firstlogic and Group 1 Software, plus several of the smaller ones, have been acquired in recent years and matching functions are now embedded within many enterprise software products. Even though there have been some intriguing new technical approaches from vendors like Netrics and Zoomix, this is a hard market to penetrate based on better technology alone. It seems that DataFlux moved into governance more in response to customer requests than due to proactive strategic planning. But even so, they have done well to recognize and seize the opportunity when it presented itself. Not everyone is quite so responsive. The question now is whether other data quality vendors will take a similar approach or this will be a long-term point of differentiation for DataFlux.

Thursday, November 29, 2007

Low Cost CDI from Infosolve, Pentaho and StrikeIron

As I’ve mentioned in a couple of previous posts, QlikView doesn’t have the built-in matching functions needed for customer data integration (CDI). This has left me looking for other ways to provide that service, preferably at a low cost. The problem is that the major CDI products like Harte-Hanks Trillium, DataMentors DataFuse and SAS DataFlux are fairly expensive.

One intriguing alternative is Infosolve Technologies. Looking at the Infosolve Web site, it’s clear they offer something relevant, since two flagship products are ‘OpenDQ’ and ‘OpenCDI’ and their tag line is ‘The Power of Zero Based Data Solutions’. But I couldn't figure out exactly what they were selling since they stress that there are ‘never any licenses, hardware requirements or term contracts’. So I broke down and asked them.

It turns out that Infosolve is a consulting firm that uses free open source technology, specifically the Pentaho platform for data integration and business intelligence. A Certified Development partner of Pentaho, Infosolve has developed its own data quality and CDI components on the platform and simply sells the consulting needed to deploy it. Interesting.

Infosolve Vice President Subbu Manchiraju and Director of Alliances Richard Romanik spent some time going over the details and gave me a brief demonstration of the platform. Basically, Pentaho lets users build graphical workflows that link components for data extracts, transformation, profiling, matching, enhancement, and reporting. It looked every bit as good as similar commercial products.

Two particular points were worth noting:

- the actual matching approach itself seems acceptable. Users build rules that specify which fields to compare, the methods used to measure similarity, and similarity scores required for each field. This is less sophisticated than the best commercial products, but field comparisons are probably adequate for most situations. Although setting up and tuning such rules can be time-consuming, Infosolve told me they can build a typical set of match routines in about half a day. More experienced or adventurous users could even do it for themselves; the user interface makes the mechanics very simple. A half-day of consulting might cost $1,000, which is not bad at all when you consider that the software itself is free. The price for a full implementation would be higher since it would involve additional consulting to set up data extracts, standardization, enhancement and other processes, but cost should still be very reasonable. You’d probably need as much consulting with other CDI systems where you'd pay for the software too.

- data verification and enhancement is done by calls to StrikeIron, which provides a slew of on-demand data services. StrikeIron is worth knowing about in its own right: it lets users access Web services including global address verification and corrections; consumer and business data lookups using D&B and Gale Group data; telephone verification, appends and reverse appends; geocoding and distance calculations; Mapquest mapping and directions; name/address parsing; sales tax lookups; local weather forecasts; securities prices; real-time fraud detection; and message delivery for text (SMS) and voice (IVR). Everything is priced on a per use basis. This opens up all sorts of interesting possibilities.

The Infosolve software can be installed on any platform that can run Java, which is just about everything. Users can also run it within the Sun Grid utility network, which has a pay-as-you-go business model of $1 per CPU hour.

I’m a bit concerned about speed with Infosolve: the company said it takes 8 to 12 hours to run a million record match on a typical PC. But that assumes you compare every record against every other record, which usually isn’t necessary. Of course, where smaller volumes are concerned, this is not an issue.

Bottom line: Infosolve and Pentaho may not meet the most extreme CDI requirements, but they could be a very attractive option when low cost and quick deployment are essential. I’ll certainly keep them in mind for my own clients.

Wednesday, June 20, 2007

Using Lifetime Value to Measure the Value of Data Quality

As readers of this blog are aware, I’ve reluctantly backed away from arguing that lifetime value should be the central metric for business management. I still think it should, but haven’t found managers ready to agree.

But even if LTV isn’t the primary metric, it can still provide a powerful analytical tool. Consider, for example, data quality. One of the challenges facing a data quality initiative is how to justify the expense. Lifetime value provides a framework for doing just that.

The method is pretty straightforward: break lifetime value in its components and quantify the impact of a proposed change on whichever components will be affected. Roll this up to business value, and there you have it.

Specifically, such a breakdown would look like this:

Business value = sum of future cash flows = number of customers x lifetime value per customer

Number of customers would be further broken down into segments, with the number of customers in each segment. Many companies have a standard segmentation scheme that would apply to all analyses of this sort. Others would create custom segmentations depending on the nature of the project. Where a specific initiative such as data quality is concerned, it would make sense to isolate the customer segments affected by the initiative and just focus on them. (This may seem self-evident, but it’s easy for people to ignore the fact that only some customers will be affected, and apply estimated benefits to everybody. This gives nice big numbers but is often quite unrealistic.)

Lifetime value per customer can be calculated many ways, but a pretty common approach is to break it into three major factors:

- acquisition value, further divided into the marketing cost of acquiring a new customer, the revenue from that initial purchase, and the fulfillment costs (product, service, etc.) related to that purchase. All these values are calculated separately for each customer segment.

- future value, which is the number of active years per customers times the value per year. Years per customer can be derived from a retention rate or a more advanced approach such as a survivor curve (showing number of customers remaining at the end of each year). Value per year can be broken into the number of orders per year times the value per order , or the average mix of products times the value per product). Value per order or product can itself be broken into revenue, marketing cost and fulfillment cost.

Laid out more formally, this comes to nine key factors:

- number of customers

- acquisition marketing cost per customer
- acquisition revenue per customer
- acquisition fulfillment cost per customer

- number of years per customer
- orders per year
- revenue per order
- marketing cost per order
- fulfillment cost per order

This approach may seem a little too customer-centric: after all, many data quality initiatives relate to things like manufacturing and internal business processes (e.g., payroll processing). Well, as my grandmother would have said, feh! (Rhymes with ‘heh’, in case you’re wondering, and signifies disdain.) First of all, you can never be too customer-centric, and shame on you for even thinking otherwise. Second of all, if you need it: every business process ultimately affects a customer, even if all it does is impact overhead costs (which affect prices and profit margins). Such items are embedded in the revenue and fulfillment cost figures above.

I could easily list examples of data quality changes that would affect each of the nine factors, but, like the margin of Fermat’s book, this blog post is too small to contain them. What I will say is that many benefits come from being able to do more precise segmentation, which will impact revenue, marketing costs, and numbers of customers, years, and orders per customer. Other benefits, impacting primarily fulfillment costs (using my broad definition), will involve more efficient back-office processes such as manufacturing, service and administration.

One additional point worth noting is many of the benefits will be discontinuous. That is, data that's currently useless because of poor quality or total absence does not become slightly useful because it becomes slightly better or partially available. A major change like targeted offers based on demographics can only be justified if accurate demographic data is available for a large portion of the customer base. The value of the data therefore remains at zero until a sufficient volume is obtained: then, it suddenly jumps to something significant. Of course, there are other cases, such as avoidance of rework or duplicate mailings, where each incremental improvement in quality does bring a small but immediate reduction in cost.

Once the business value of a particular data quality effort has been calculated, it’s easy to prepare a traditional return on investment calculation. All you need to add is the cost of improvement itself.

Naturally, the real challenge here is estimating the impact of a particular improvement. There’s no shortcut to make this easy: you simply have to work through the specifics of each case. But having a standard set of factors makes it easier to identify the possible benefits and to compare alternative projects. Perhaps more important, the framework makes it easy to show how improvements will affect conventional financial measurements. These will often make sense to managers who are unfamiliar with the details of the data and processes involved. Finally, the framework and related financial measurements provide benchmarks that can later be compared with actual results to show whether the expected benefits were realized. Although such accountability can be somewhat frightening, proof of success will ultimately build credibility. This, in turn, will help future projects gain easier approval.