Showing posts with label unified customer data. Show all posts
Showing posts with label unified customer data. Show all posts

Wednesday, January 24, 2018

Simple Questions to Screen Customer Data Platform Vendors

I’ve been working for months to find a way to help marketers understand the differences between Customer Data Platform vendors. After several trial balloons and with considerable help from industry friends, I recently published a set of criteria that I think will do the job. You can see the full explanation on the CDP Institute blog. But, since this blog has its own readership I figured I’d post the basics here as well.

The primary goal is give marketers a relatively easy way to decide which CDPs are likely to meet their needs. To do this I’ve come up with a a small list of features that relate directly to working with particular data sources and supporting particular applications. The theory is that marketers know what sources and applications they need to support, even if they're not experts in the fine points of CDP technology.

In other words, read these items as meaning: if you want your CDP to support [this data type or application] then it should have [this feature].

Obviously this list covers just a tiny fraction of all possible CDP features. It’s up to marketers to dig into the details of each system to determine how well it supports their specific needs.  We have detailed lists of CDP features in the Evaluation section of the CDP Institute Library.

The final list also includes a few features that are present in all CDPs (or, more precisely, in all systems that I consider a CDP – we can’t control what vendors say about themselves). These are presented since there’s still some confusion about how CDPs differ from other types of systems.

Now that the list is set, the next step is to research which features are actually present in which vendors and publish the results. That will take a while but when it’s done I’ll certainly announce it here.

Here’s the list:

Shared CDP Features: Every CDP does all of these. Non-CDPs may or may not.
  • Retain original detail. The system stores data with all the detail provided when it was loaded. This means all details associated with purchase transactions, promotion history, Web browsing logs, changes to personal data, etc. Inputs might be physically reformatted when they’re loaded into the CDP but can be reconstructed if needed.
  • Persistent data. The system retains the input data as long as the customer chooses. (This is implied by the previous item but is listed separately to simplify comparison with non-CDP systems.)
  • Individual detail. The system can access all detailed data associated with each person. (This is also implied by the first item but is a critical difference from systems that only store and access segment tags on customer records.)
  • Vendor-neutral access. All stored data can be exposed to any external system, not only components of the vendor’s own suite. Exposing particular items might require some set-up and access is not necessarily a real time query.
  • Manage Personally Identifiable Information (PII). The system manages Personally Identifiable Information such as name, address, email, and phone number. PII is subject to privacy and security regulations that vary based on data type, location, permissions, and other factors.
Differentiating CDP Features: A CDP doesn’t have to do any of these although many do some and some do many. These are divided into three subclasses: data management, analytics, and customer engagement.

Data Management. These are features that gather, assemble, and expose the CDP data.

     Base Features. These apply to all types of data.
  • API/query access. External systems can access CDP data via an API or standard query language such as SQL. It’s just barely acceptable for a CDP to not offer this function and instead provide access through data extracts. But API or query access is much preferred and usually available. API or query access often requires some intermediate configuration, reformatting, or indexing to expose items within the CDP’s primarily data store. Those are important details that buyers must explore separately.
  • Persistent ID. The system assigns each person an internal identifier and maintains it over time despite changes or multiple versions of other identifiers, such as email address or phone number. This allows the CDP to maintain individual history over time, even when source systems might discard old identifiers. CDPs that use a persistent ID applied outside of the system do not meet this requirement.
  • Deterministic match (a.k.a. “identity stitching”). The system can store multiple identifiers known to belong to the same person and link them to a shared ID (usually the persistent ID). This enables the system to connect identifiers indirectly: for example, if an email linked to an account is opened on a particular device, subsequent activity on that device can also be linked to the account.
  • Probabilistic match (a.k.a. “cross device match”). The system can apply statistical methods and rules to identify multiple devices used by the same person, such as computers, tablets, smart phones, and home appliances. While many CDPs rely on third party services for this sort of matching, this item refers only to matching done by the CDP itself.
     Unstructured and Semi-Structured Data. This refers to loading data from unstructured or semi-structured sources such as Web logs, social media comments, voice, video, or mages. These are typically managed with “big data” technologies such as Hadoop. Nearly all CDPs use some version of this technology but it’s only essential if clients have unstructured or semi-structured sources and/or very high data volumes. Some CDPs handle very high data volumes in structured databases such as Amazon Redshift.
  • JSON load. The system can accept and store data through JSON feeds without the user specifying in advance the specific attributes that will be included. Additional configuration may later be required to access this data. There are some alternatives to JSON that offer similar capabilities.
  • Schema-free data store. The system uses a data store that does not require advance specification of the elements to be stored. Examples include Hadoop, Cassanda, MongoDB, and Neo4J.
     Web Site. This refers to interactions with the company’s own Web site, whether on a desktop computer or mobile device.
  • Javascript tag. The system provides a Javascript tag that can be loaded into the client’s Web site and used to capture data about customer behaviors. Some CDP vendors provide full tag management systems but this is not a requirement for this item. This item does require that data captured by the Javascript tag can be associated with a customer record in the CDP database. This is usually done with a Web tracking cookie but sometimes through other methods.
  • Cookie management. The system can deploy and maintain Web browser cookies associated with the client’s own Web site. The cookies can be linked to customer records in the CDP database.
     Mobile Apps. This refers to interactions with mobile apps created by the company.
  • SDK load. The system offers a Software Development Kit (SDK) that can load data from a mobile app into the CDP database. It must be able to associate the data with individual customers in the CDP database. This is usually done through an app ID. Other SDK features such as message delivery are not a requirement for this item.
     Display Ads. This refers to interactions through display advertising networks, including social media networks.
  • Audience API. The system has an API that can send customer lists from the CDP to systems that will use them as advertising audiences. The receiving systems might be Data Management Platforms, Demand Side Platforms, advertising exchanges, social media publishers, or others. Ability to receive information back from the advertising systems is not a requirement for this item.
  • Cookie synch. The CDP can match its own cookie IDs with third party cookie IDs to allow the marketer to enrich profiles with external data or reach users through advertising networks.
     Offline. This refers to interactions managed through offline sources such as direct mail and retail stores, where the customer’s primary identifier is name and postal address.
  • Postal Address. The system can clean, standardize, verify, and otherwise work with postal addresses. This processing is reduces inconsistencies and makes matching more effective. Systems meet this requirement so long as the address processing is built into system process flows, even if they rely on third party software. Systems that send records to external systems in a batch process do not meet this requirement.
  • Name/Address Match. The system can find matches between different postal name/address records despite variations in spelling, missing data elements, and similar differences. As with postal processing, systems can meet this requirement with third party matching software so long as the software is embedded in their processing flows.
     Business to Business. This refers to companies that sell to other businesses rather than to consumers.
  • Account-level data. The system can maintain separate customer records for accounts (i.e., businesses) and for individuals within those accounts. This means account information is stored and updated separately from individual information. It also means that selections, campaigns, reports, analyses, and other system activities can combine data from both levels.
  • Lead to Account Match. The system can determine which individuals should be associated with which account records, using information such as company name, address, email domain, and telephone number. This excludes processing done by sending batch files to external vendors.
Analytics. These are applications that use the CDP data but don’t extend to selecting messages, which is the province of customer engagement.
  • Segmentation. The system lets non-technical users define customer segments and automatically send segment member information to external systems on a user-defined schedule. Ideally, all data would be available to use in the segment definitions and to include in the extract files. In practice, some configuration may be needed to expose particular elements. Systems meet this requirement regardless of whether segments are defined manually or discovered by automated processes such as cluster analysis.
  • Incremental attribution. The system has algorithms to estimate the incremental impact of different marketing activities on specified outcomes such as a purchase or conversion. Attribution is a specialized analytical process that relies on the unified customer data assembled by the CDP. Algorithms vary greatly. To qualify for this item, the algorithm must estimate the contribution of different marketing contacts on the final result. That is, fixed approaches such as “first touch” or “U-shaped distribution” are not included.
  • Automated predictive. The system can generate, deploy, and refresh predictive models without involvement of a technical user such as a data scientist or statistician. This usually employs some form of machine learning. There are many different types of automated predictive; systems meet this requirement if they have any of them.
 Engagement. This refers to applications that select messages for individual customers. It does not include content delivery, which is typically handled outside of the CDP.
  • Content selection. The system can select appropriate marketing or editorial content for individual customers in the current situation, based on the data it stores about them, other information, and user instructions. The instructions may employ fixed rules, predictive models, or a combination. Selections may be made as part of a batch process.
  • Multi-step campaigns. The system can select a series of marketing messages for individual customers over time, based on data and user instructions. The message sequence is defined in advance but may change or be terminated depending on customer behaviors as the sequence is executed.
  • Real-time interactions. The system can select appropriate marketing or editorial content for individual customers during a real-time interaction. This requires accepting input about the customer from a customer-facing system, finding that customer’s data within the CDP, selecting appropriate content, and sending the results back to the customer-facing system for delivery. The results might include the actual message or instructions that enable the customer-facing system to generate the message.

Wednesday, January 18, 2017

Customer Data Platform Industry Profile: A Look Inside the Numbers

My snarky twin at the Customer Data Platform Institute just published a new report on the CDP industry. Since few industry vendors release financial or business details, the report relies on public sources including Owler for revenue estimates, Crunchbase for funding history, and LinkedIn for employee counts. Most vendors did provide client counts, and several privately shared other information where the public data was clearly wrong. You can download the report here. I'll wait while you do that.  (Sound of fingers tapping.)


Okay, you've downloaded it, right?  Good.

As you see, the report only presents figures for the industry as a whole. We feel those are reasonably accurate but that data for individual vendors are too unreliable to show separately. That may sound illogical but bear in mind that figures for the larger vendors are more reliable, so many errors that are significant for individual small vendors don’t materially change the total. Also remember that some vendors provided information in confidence and we made estimates of our own for some others.

I do feel I can safely publish statistics for three groups within the industry.  This gives some additional insight without exposing any proprietary or misleading vendor data.  The groups are based on each vendor's original business.  They are:

  • Tag managers. This may seem an unlikely starting point, but it actually makes sense.  Tag management was originally about collecting data once (when a Web page loaded) and then sharing it with other systems that would otherwise have their own tags. This gave the Web site owner more control over what went where and reduced the number of tags on each page.  The data sharing was similar to what happens in integration platforms/data hubs  like Jitterbit and Zapier. So tag managers were always about data distribution.  To become true CDPs, the tag vendors had to ingest data from additional sources and send the data to a persistent database. Ingesting new sources can be challenging but vendors could grow incrementally by choosing which sources to accept.  Feeding a persistent data is basically just adding a new destination for data sharing.  So the transition to CDP offered a reasonable path to escape being a commodity tag manager.

  • Campaign managers. I’m using this term loosely to include companies that offered any sort of marketing message selection. It includes systems that do email, Web site messages, mobile app messages, and omnichannel campaigns. These vendors all started out as CDPs in the sense that they always built unified customer databases. Among other things, this meant that most included reasonably robust cross-channel identity resolution. These vendors didn’t necessarily start by sharing their database with other systems.  But they do it now or I wouldn't consider them a CDP.

  • Data assembly systems. This is a bit of a catch-all category but almost every system in this group was designed primarily to create a customer database that would be accessible to other systems.  Intended uses included analytics, marketing execution, or both. (I say "almost" because the group includes two systems that built databases primarily to support their own attribution services.) There’s more variety within this group than the other two.  But many vendors provide advanced identity resolution and all are strong at providing external access.
Here are key statistics for each group.



original purpose:

vendors

funding

2016 revenue

customer count

revenue / customer

employee count

revenue / employee

Tag Management

6

$356 million

$118  million

13,500

$9,000

840

$141,000

Campaign Management

8

$106 million

$108 million

1,000

$108,000

520

$207,000

Data Assembly
13
$246 million
$100 million
3,000
$34,000
920
$108,000

Here are some observations:

  • Revenue is split about evenly among the three groups. That’s a bit surprising because tag management is an older and more established category than the others, so you might have expected it to have more revenue.  Vendors in the other categories do tend to be newer, smaller, and growing more quickly.

  • Tag management vendors have many more customers and earn less revenue per customer. This largely reflects the original tag management products, which are sold to many non-enterprise customers. But the tag management vendors also have hundreds of enterprise clients.  Many of those clients are building the large-scale customer databases we expect to call a CDP.  The tag management group also includes a couple of vendors who specialize in building CDPs for smaller companies. These are not as expensive as the enterprise installations.  The campaign management vendors average about $100,000 per customer, which is what you'd expect for an enterprise CDP.  Revenue per customer is just $34,000 for the data assembly vendors, but that's largely due to one vendor with 2,000 clients.  Without that vendor, revenue per customer figures for the data assembly group would be $88,000. Backing out non-CDP clients is why the report puts the number of CDP customers for the entire industry at 2,500.
  • Revenue per employee is generally in line with what we expect to see at growing Software as a Service companies. The standout here is the campaign management group, which has an impressively high figure of $207,000 per employee.  This suggests the campaign managers have a high value-added business.  The relatively low amount of external funding is more evidence that campaign managers throw off considerable cash from their own operations. The much lower revenue per employee for the data assembly companies, $108,000, is more typical of new SaaS ventures.  Indeed, several of these are just starting to earn revenue from their first clients.  These data assembly companies have attracted considerably more funding than the campaign managers, giving them a cushion to invest in growth. (If you’re wondering about that company with 2,000 clients, its revenue per employee is similar to others in its group.)

There are other nuances to consider in assessing these figures. For example, several vendors do business through agencies, which makes it harder to count clients and to compare revenue per client.  But the over-all picture that emerges is a healthy industry that is already attracting substantial revenue and funding.

The report projects a 50% annual growth rate, which yields an estimated $1 billion revenue for 2019.  The projection is based on public and private reported growth rates, which actually averaged much higher than 50% on a revenue-weighted basis.  The report used 50% to be conservative.  While past performance doesn't guarantee future growth, I think CDP revenues will if anything accelerate because most marketers still don't realize what a CDP can do for them.  As more of them get the message, CDP adoption should skyrocket. So the future is bright indeed.

Wednesday, November 30, 2016

3 Insights to Help Build Your Unified Customer Database

The Customer Data Platform Institute (which is run by Raab Associates) on Monday published results of a survey we conducted in cooperation with MarTech Advisor. The goal was to assess the current state of customer data unification and, more important, to start exploring management practices that help companies create the rare-but-coveted single customer view.

You can download the full survey report here (registration required) and I’ve already written some analysis on the Institute blog . But it’s a rich set of data so this post will highlight some other helpful insights.

1. All central customer databases are not equal.

We asked several different questions whose answers depended in part on whether the respondent had a unified customer database. The percentage who said they did ranged from 14% to 72%:


I should stress that these answers all came from the same people and we only analyzed responses with answers to all questions.  And, although we didn’t test their mental states, I doubt a significant fraction had multiple personality disorders. One lesson is that the exact question really matters, which makes comparing answers across different surveys quite unreliable. But the more interesting insight is there are real differences in the degree of integration involved with sharing customer data.

You’ll notice the question with the fewest positive answers – “many systems connected through a shared customer database” describes a high level of integration.  It’s not just that data is loaded into a central database, but that systems are actually connected to a shared central database. Since context clearly matters, here is the actual question and other available answers:

 The other questions set a lower bar, referring to a “unified customer database” (33%), “central database (42%) and "central customer database” (57%). Those answers could include systems where data is copied into a central database but then used only for analysis. That is, they don’t imply connections or sharing with operational customer-facing systems. They also could describe situations where one primary system has all the data and thus functions as a central or unified database.

The 72% question covered an even broader set of possibilities because it only described how customer data is combined, not where those combinations take place. That is, the combinations could be happening in operational systems that share data directly: no central database is required or even implied.  Here are the exact options:


The same range of possibilities is reflected in answers about how people would use a single customer view. The most common answers are personalization and customer insights.  Those require little or no integration between operational systems and the central database, since personalization can easily be supported by periodically synchronizing a few data elements. It’s telling that consistent treatments ranks almost dead last – even though consistent experiences are often cited as the reason a central database is urgently required.


This array of options to describe the central customer database suggests a maturity model or deployment sequence.  It would start with limited unification by sharing data directly between systems (the most common approach, based on the stack question shown above), progress to a central database that assembles the data but doesn’t share it with the operational systems, and ultimately achieve the perfect bliss of unity, which, in martech terms, means all operational systems are using the shared database to execute customer interactions.  Purists might be troubled by these shades of gray, but they offer a practical path to salvation. In any case, it’s certainly important to keep these degrees in mind and clarify what anyone means when they talk about shared customer data or that single customer view.

2. You must have faith.

Hmm, a religious theme seems to be emerging.  I hadn’t intended that but maybe it’s appropriate. In any event, I’ve long argued that the real reason technologies like marketing automation and predictive modeling don’t get adopted more quickly are not the practical obstacles or lack of proven value, but lack of belief among managers that they are worthwhile. This doesn’t show up in surveys, which usually show things like budget, organization, and technology as the main obstacles. My logic has been that those are basically excuses: people would find the resources and overcome the organizational barriers if they felt the project were important enough.  So citing budgets and organizational constraints really means they see better uses for their limited resources.

The survey data supports my view nicely. Looking at everyone’s answers to a question about obstacles, the answers are rather muddled: budget is indeed the most commonly cited obstacle (41%), followed closely by the technical barrier of extracting data from source systems (39%). Then there’s a virtual tie among organizational roadblocks (31%), other priorities in IT (29%), other priorities in marketing (29%) and systems can’t use (29%). Not much of a pattern there.

But when you divide the respondents based on whether they think single customer view is important for over-all marketing success, a stark division emerges.  Budget and organization are the top two obstacles for people who don’t think the unified view is needed, while having systems that can extract and use the data are top two obstacles for people who do think it’s necessary for success. In other words, the people committed to unified data are focused on practical obstacles, while those who don’t are using the same objections they apply to everything else.


Not surprisingly, people who classify SCV as extremely important are more likely to actually have a database in place than people who consider it just very important, who in turn have more databases than people who consider it even less important or not important at all.  (In case you're wondering, each group accounts for roughly one-third of the total.)

The same split applies to what people would consider helpful in build in building a single customer view: people who consider the single view important are most interested in best practices, case studies, and planning assumptions – i.e., building a business case.  Those who think it’s unimportant ask for product information, vendor lists, and pricing. I find this particular split a bit puzzling, since you’d think people who don’t much care about a unified database would be least interested in the details of building one. A cynic might say they’re looking for excuses (cost is too high) but maybe they’re actually trying to find an easy solution so they can avoid a major investment.

Jumping ahead just a bit, the idea that SCV doubters are less engaged than believers also shows up in at the management tools they use.  People who rated SCV as extremely important were much more likely to use all the tools we asked about. Interestingly, the biggest gap is in use of value metrics. This could be read to mean that people become believers after they measure the value of a central database, or that people set up measurements after they decide they need to prove their beliefs. My theology is pretty rusty but surely there’s a standard debate about whether faith or action comes first.

Regardless of the exact reasons for the different attitudes, the fundamental insight here is that people who consider a single view important act quite differently from people who don’t. This means that if you’re trying to sell a customer database, either in your own company or as a vendor, you need to understand who falls into which category and address them in appropriate terms. And I guess a little prayer never hurt.

3. Tools matter.

We’ve already seen that believers have more databases and have more tools, so you won’t be surprised that using more tools correlates directly with having or planning a database.


Let's introduce the tools formally.  Here are the exact definitions we used and the percentage of people who said each was present in their organization:


Of course, the really interesting question isn’t which tools are most popular but which actually contribute (or at least correlate) with deploying a database. We looked at tool use for three groups: people with a database, people planning a database, and people with no such plans. 

Over all, results for the different tools were pretty similar: people who used each tool were much more likely to have a database and somewhat more likely to plan to build one. The pattern is a bit jumbled for Centers of Excellence and technology standards, but the numbers are small so the differences may not be significant. But it's still worth noting that Centers of Excellence are really tools to diffuse expertise in using marketing technology and don’t have too much to do with actually creating a customer database.

If you’re looking for a dog that didn’t bark, you might have expected companies using agile to be exceptionally likely to either have a database or be planning one. All quiet on that front: the numbers for agile look like numbers for long term planning and value metrics, adjusting for relative popularity. So agile is helpful but not a magic bullet.

What have we learned here? 

Clearly, we've learned that management tools are important and that long term planning in particular both the most common and the best predictor of success.

We also found that tools aren’t enough: managers need to be convinced that a unified customer view is important before they’ll invest in a database or tools to build it.

And, going back to the beginning, we saw that there are many forms of unified data, varying in how data is shared, where it’s stored, how it’s unified, and how it’s used. While it’s easy enough to assume that tight, real-time integration is needed to provide unified omni-channel customer experiences, many marketers would be satisfied with much less. I’d personally hope to see more but, as every good missionary knows, people move towards enlightenment in many small steps.

Friday, October 28, 2016

Singing the Customer Data Platform Blues: Who's to Blame for Disjointed Customer Data?

I’m in the midst of collating data from 150 published surveys about marketing technology, a project that is fascinating and stupefying at the same time. A theme related to marketing data seems to be emerging that I didn’t expect and many marketers won’t necessarily be happy to hear.

Most surveys present a familiar tune: many marketers want unified customer data but few have it. This excerpt from an especially fine study by Econsultancy makes the case clearly although plenty of other studies show something similar.



So far so good. The gap is music to my ears, since helping marketers fill it keeps consultants like me in the business. But it inevitably raises the question of why the gap exists.

The conventional answer is it’s a technology problem. Indeed, this Experian survey makes exactly that point: the top barriers are all technology related.



And, comfortingly, marketers can sing their same old song of blaming IT for failing to deliver what they need.  For example, even though 61% of companies in this Forbes Insights survey had a central database of some sort, only 14% had fully unified, accessible data.



But something sounds a little funny. After all, doesn’t marketing now control its own fate? In this Ascend2 report, 61% of the marketing departments said they were primarily responsible for marketing data and nearly all the other marketers said they shared responsibility.



Now we hear that quavering note of uncertainty: maybe it’s marketing’s own fault? That’s something I didn’t expect. And the data seems to support it. For example, a study from Black Ink ROI found that the top barrier to success was better analytics (which implicitly requires better data) and explicitly listed data access as the third-ranked barrier.


But – and here’s the grand finale – the same study found that data integration software ranked sixth on the marketers’ shopping lists. In other words, even though marketers knew they needed better data, they weren’t planning to spend money to make it happen. That’s a sour chord indeed.



But the song isn't over.  If we listen closely, we can barely make out one final chorus: marketers won’t invest in data management technology because they don’t have the skills to use it. Or that’s what this survey from Falcon.io seems to suggest.



In its own way, that’s an upbeat ending. Expertise can be acquired through training or hiring outside experts (or possibly even mending some fences with IT). Better tools, like Customer Data Platforms, help by reducing the expertise needed. So while marketers aren't strutting towards a complete customer view with a triumphal Sousa march, there’s no need for a funeral dirge quite yet.