lunes, 6 de diciembre de 2010

Users Trends (Business analysts)

On the other side of the equation, power users require MAD capabilities 20 to 40% of the time. The bulk of their time is spent using tools designed to handle a variety of analytical tasks, including report authoring tools, spreadsheet-based modeling tools, sophisticated OLAP and visual design tools, and predictive modeling and data mining tools.

Times have never been better for power users. Their desktop computers contain more processing power and can hold more data than ever before. Today, there are more tools designed to help power users exploit these computing resources to analyze information. Many cost less than $1,000 for a single user or can be downloaded from the Internet. “Power users have more power today than ever to perform deep analytics,” .

Despite the plentiful options, many power users are bereft of optimal analytical tools. Either they restrict themselves to spreadsheets and desktop databases, or that’s all their organization will give them. Most homeowners wouldn’t hire a carpenter with just one or two tools in his toolbox; they want a carpenter whose toolbox contains tools for every type of carpentry task imaginable. In the same way, organizations need to empower power users with a multitude of tools and technologies to make them more productive as analysts. If implemented correctly, the technology can liberate analysts to gather, analyze, and present data quickly and efficiently without undermining enterprise IT standards governing data, semantics, and tools.

Four types. Power users are a diverse group who perform a variety of analytical tasks. I’ve divided power users into four types:

1. Business analysts. Data- and process-savvy business users who use data to identify trends, solve problems, and devise plans.

2. Super users. Technically savvy departmental business users who create ad hoc reports on behalf of their colleagues.

3. Analytical modelers. Business analysts who create statistical and data mining models that quantify relationships and can be used to predict future behavior or conditions.

4. IT report developers. IT developers, analysts, or administrators who create complex reports and train and support super users.

According to our survey, most organizations have all four types of power users, although only 51% have analytical modelers.

 

image

 

BUSINESS ANALYSTS. Business analysts sit at the intersection of data, process, and strategy, and they play a significant role in helping the business solve problems, devise plans, and exploit opportunities. Their titles include “business analyst,” “financial analyst,” “marketing specialist,” and “operations research analyst.” Executives view them as critical advisors who keep them grounded in reality (data) and help them bolster arguments for courses of action.

Business analysts perform three major tasks:
1. Gather data. Analysts explore the characteristics of various data sets, extract desired data, and transform the extracted data into a standard format for analysis.

2. Analyze data. Analysts examine data sets in an iterative fashion—essentially “playing with the data”—to identify trends or root causes. Analysts will visualize, aggregate, filter, sort, rank,
calculate, drill, pivot, model, and add or delete columns, among other things.

3. Present data. Analysts deliver the results of their analysis to others in a standard format, such as a report, presentation, spreadsheet, PDF document, or dashboard.


Today, business analysts spend an inordinate amount of time on steps 1 and 3 and not enough time on step 2, which is what they were hired to do. Unfortunately, due to the sorry state of data in most organizations, they have become human data warehouses. TDWI estimates that business analysts spend an average of two days every week gathering and formatting data instead of analyzing it, costing organizations an average of $780,000 a year.

According some survey, most business analysts use spreadsheets to access, analyze, and present data, followed by BI reporting and analysis tools. However, in most cases, the analysts use BI tools as glorified extract tools to grab data warehouse data and dump it into a spreadsheet or desktop database, where they normalize the data and then analyze it. The next most popular tool is SQL, which analysts use to access operational and other sources so they can dump the data into spreadsheets or desktop databases (which rank number five on the list, following OLAP tools).

 

image

 

To improve the productivity and effectiveness of business analysts, organizations should continue to expand the breadth and depth of their data warehouses, which will reduce the number of data sources that analysts need to access directly. They should also equip analysts with better analytical tools that operate the way they do. These types of tools include speed-of-thought analysis (i.e., subsecond responses to all actions) and better visualizations to spot outliers and trends more quickly.

lunes, 25 de octubre de 2010

Traditional BI Systems A re N ot Des igned For A gility


The traditional approach to business intelligence (BI) has reached its limits. Over the past 20 years we have developed a set of procedures and technologies that allow us to aggregate,
cleanse, prepare, and report on enterprise data. While there have been incremental improvements in efficiency and speed over that time, the fundamental approach to BI has not changed much. During that same period, however, businesses and their information needs have changed dramatically. Enterprises have transformed themselves from rather isolated entities, often with an inward focus on operational efficiency, to players in geographically dispersed, multi-partied ecosystems that need to have as much understanding of their customer’s interests and partners’ operations as they do their own operational efficiency. The information they collect, transport, and analyze has changed accordingly and, more importantly, so has the speed with which they need to act.

Today agility is the goal of most organizations regardless of their industry and it is a top goal of business leaders whether they are in sales, marketing, manufacturing, engineering, customer care, or service delivery. Businesses are competing not just on efficiency, but on their ability to sense market conditions and quickly respond. That is exactly where BI and the business needs have diverged.

In many organizations, BI systems are used almost exclusively to generate standard monthly or quarterly reports. These reports deliver great value – certainly most organizations could not survive without them. But traditional BI systems require an expensive and time-consuming process to identify all the answers the business users ultimately want, build a data model that captures that information and unify data from disparate sources into that data model. Often, multiple cycles of this process are required before the needs of the business users are fully addressed. More than anything, a traditional BI approach is designed to structure data, ensure consistency across systems, and efficiently handle large amounts of structured data – goals that seemingly run counter to the need for agility.

jueves, 7 de octubre de 2010

Q&A on SaaS BI Technology


SaaS is a service model that sits on top of cloud computing environments and is designed to leverage the power of the cloud infrastructure. The software applications are accessed via a client such as a Web browser. They are managed and maintained by the SaaS vendor company removing much of the administration costs generally associated with on-premises software solutions.

SaaS and cloud computing are dependent on one another to bring their unique value to the market.


What types of projects are most often addressed with SaaS BI? What BI projects aren't good candidates for this technology?

The types of projects best suited for SaaS business intelligence are evolving quickly. Several years ago, the most basic functions of BI were the only ones best suited for the SaaS environment. Recently, SaaS vendors have taken great strides to deliver sophisticated feature sets and are now branching off into corporate performance management suites and even on-demand predictive analytics. Reporting and analysis still lead the way as the most heavily adopted feature sets within SaaS solutions, but as the technology continues to evolve, so will the demands of the SaaS customers.

Because of the obvious data integration challenges presented by SaaS Bi applications, real-time data is still difficult to leverage in a SaaS model. Leading integration vendors are starting to deliver new solutions that have greatly reduce data access time. As these innovations continue, real-time business intelligence will become easier.


One of the benefits SaaS promises is that it reduces an enterprise's IT costs. Is SaaS really less costly than traditional solutions?

In many cases, SaaS BI is less expensive than traditional on-premises software. Recent research by Enterprise Management Associates (EMA) shows that 76 percent of the organizations implementing SaaS have realized significant ROI. Upfront capital costs of hardware, along with the less expensive operating costs of SaaS, have contributed to these savings. There are exceptions, especially when the number of user licenses is high. Many companies have moved to a total cost of ownership (TCO) model when analyzing the impact of SaaS versus on-premises solutions. Many variables are involved in these scenarios and determine which direction is financially best for a company.


What other benefits can BI professionals expect?

There are many benefits to SaaS BI from a monetary standpoint. Both the capital and operating expenses of IT are reduced for most SaaS implementations. This upfront cost savings can translate into reduced risk for companies that are trying to experiment or deliver proof-of-concept (POC) projects. The elastic qualities of SaaS are a perfect fit for companies that experience seasonal highs and lows, as they don’t need to purchase additional hardware to handle the changes. Some studies have shown that SaaS BI solutions have fostered greater adoption among business users, adding to a more diverse business intelligence community.


You've discussed some compelling benefits. Are there any drawbacks to SaaS? Given that data is now "off site," isn't security a problem?

Security has always been a hurdle for SaaS. Customers are vigilant about securing their data as well as governance, regulatory, and compliance issues that surround it. The vendor community recognized this challenge early on and has addressed it on two fronts. The first is third-party auditing and certification. Leading SaaS BI vendors are SAS-70 certified and some have also attained Systrust certification. Both certifications relate to how the company controls client data and the systems and processes surrounding their data infrastructure.

The second front is innovation around data integration. Many of the vendors are finding ways to keep the data where it is while still leveraging the power of SaaS and cloud computing platforms. Data virtualization firms have also entered the market, providing trusted data federated environments that reduce the security risks of off-premise systems. In the end, the most reliable security feature for SaaS vendors is their own desire to prosper. EMA research has shown that 83 percent of the respondents would be unlikely to work with a firm that has had a security event. SaaS vendors know this and take every precaution to secure customer data and the applications they run on.


Is SaaS a real alternative for enterprise-sized companies? If so, what are the challenges to success?

Small to mid-size companies make up the largest portion of SaaS BI customers. Enterprise-sized companies have been slower to adopt SaaS BI. Large companies are often not early adopters and take a wait-and-see perspective. This, coupled with the need for a greater level of sophistication around feature sets, made it difficult for these firms to see SaaS as a viable alternative to on premise solutions early on. Enterprise-level successes such as SalesForce.com have shown large firms that SaaS can fit their needs and expectations, causing them to look hard at what the SaaS BI market has to offer. There is now a thriving community of firms big and small delivering SaaS BI to enterprise clients such as AON Insurance, ACNeilson, DHL, and Citrix. Both pure-play upstarts and the biggest BI vendors in the world are now offering on-demand SaaS BI products and applications.


What best practices can you offer to overcome these challenges and for successfully managing a SaaS BI project?

Companies looking into SaaS BI solutions need to consider many things, including security, licensing, TCO, feature sets, training, and service management. I highly recommend that before jumping into the deep end of the pool, a customer should scope out a proof-of-concept project that closely aligns with their overall business intelligence needs.

Like any BI project stakeholders, IT staff and the line-of-business users must take part in the process. It’s important that the customer understands both the flexibility and the restrictions of a SaaS application. Customization in a SaaS environment can be costly or impractical. Support and service are also important in a SaaS relationship. Service events should be built into the POC to assure that the needs of the clients are met.

lunes, 27 de septiembre de 2010

Designing a Metrics Dashboard for the Sales Organization

The primary objective of the dashboard creation process is to identify and implement key performance measures and indicators that will enable managers to quickly and effectively manage the sales organization. This can be accomplished through selecting metrics that support sales objectives, strategy and goals. Some of the benefits that will result from implementing the dashboard include:


• Gain a deeper understanding of the drivers of sales productivity
• Identify where management action is required to improve sales productivity and effectiveness
• Develop a common vehicle for monitoring and improving performance
• Understand sales performance from a variety of perspectives
• Build consensus on key performance measures and drivers
• Clarify accountability around specific measures
• Enable performance benchmarking with competitors and best-in-class companies


Approach
Corporate vision guides the development of an organization’s sales objectives, strategy and tactical goals. Metrics are in turn driven by sales strategy and goals. At the tactical level, metrics serve as the primary vehicle for managing performance within the organization. Targets are set for each metric, performance is monitored and interpreted to provide timely feedback and corrective actions are initiated.

But which metrics should we choose? The sheer abundance of metrics creates a situation in which it may be difficult to properly identify metrics that make the most sense. In answering this question, the first step is to create a framework in which all the available metrics may be organized and prioritized. This framework consists in two dimensions; first, a corporate perspectives dimension and secondly a sales performance dimension. The corporate approach takes a 360 degree view of the organization from five distinct perspectives: customers, employees, partners, investors and internal processes. This approach is typically utilized in the so called “Balanced Scorecard” approach.

Each of the corporate perspectives should be examined and appropriate individuals identified to provide a list of metrics. In addition to the corporate perspective, a sales performance dimension must also be included. This breaks sales performance into four elements: readiness, productivity, efficiency and effectiveness.

Each of the corporate perspectives should be examined and appropriate individuals identified to provide a list of metrics. In addition to the corporate perspective, a sales performance dimension must also be included. This breaks sales performance into four elements: readiness, productivity, efficiency and effectiveness.


The key to the metrics identification process consists in both fact-finding and identifying metrics as well as categorizing metrics according to the above two dimensions, corporate perspective and sales performance. This basically involves the creation of a matrix with these two axes which then may be populated with metrics collected through the fact-finding process.


Dashboard Design Process
The dashboard design process consists in metric selection, design and implementation. Each of these steps involve some basic principles outlined below.

Metric Selection

• Supports stated objectives, strategies and goals
• Can be directly impacted by sales management
• Can be measured in a cost effective and timely fashion
• Reflects one of the four key dimensions of sales performance (readiness, productivity, efficiency and effectiveness)
• Enables performance benchmarking with industry competitors and best-in-class companies


Dashboard Design Principles

• Reflects senior management priorities
• Balances internal and external metrics
• Includes measures of past performance and indicators of future performance
• Minimizes the number of metrics in order to facilitate management interpretation

The actual design process is outlined below along with the detailed steps involved.

1. Metric Selection

Identify existing and potential metrics by corporate performance perspective (interview process)
• Categorize metrics into four dimensions of sales performance (efficiency, effectiveness, productivity and readiness) and eliminate unclassifiable metrics
• Create preliminary scorecard matrix that combines business perspectives with sales performance dimensions
• Review scorecard matrix for completeness and add metrics based on experience


2. Dashboard Design

• Eliminate metrics that cannot be measured or are too costly to measure
• Eliminate metrics that cannot be significantly impacted by sales management
• Prioritize metrics based on alignment with stated strategy and goals
• Select top metric per cell in scorecard matrix based on alternative approaches
• Evaluate alternative scorecards and select most appropriate metrics


3. Implementation

• Assign metric accountability
• Determine performance targets
• Obtain available benchmark data
• Determine monitoring, interpretation and feedback procedures and guidelines
• Develop corrective action review process


Metrics Matrix
Design To facilitate the dashboard design process, a matrix tool may be created to help classify the various metrics uncovered in the fact finding process. Because each metric can be understood in terms of sales performance as well as a business perspective, a metrics matrix can be created that combines the business perspectives along the horizontal axis with sales performance dimensions along the vertical axis. Each metric is placed in the matrix based on its most appropriate classification with respect to these dimensions. This tool has the following benefits:

• Creates a framework around the metrics selection process
• Balances business perspectives and sales performance views
• Provides a systematic approach
• Facilitates prioritization
• Allows identification of particular areas of emphasis
• Highlights areas with no metric coverage


Criteria for Eliminating Metrics
Eliminate metrics that cannot be measured or would be too costly to measure
• Partner coverage
• Amount of effort exerted on business approvals

Eliminate metrics that cannot be directly impacted by the sales organization
• Customer’s growth rates
• Customer profitability
• Partner satisfaction
• Number of deals involving per partner
• Share of partner revenue by platform
• Partner’s profit margin
• Partner churn
• Rate of technology transfer
• Number of certified consultants
• Number of certified partners

Prioritization Decision Rules
Each cell in the metrics matrix may contain many metrics and, as a result, must be prioritized. Some basic rules to follow in that process are as follows:

• Alignment with stated strategy and goals – Use metrics that align with strategy or show alignment with strategy the organization
• Frequency and intensity of emphasis during fact-finding – Use metrics that different corporate perspectives emphasize
• Experience – Use metrics that experience shows are important to measure
• Availability of benchmark data – Use metrics for which benchmarks exist


Preliminary Dashboard
After the completion of the matrix a preliminary matrix may be created that graphically represents the top metrics from each cell. Feedback from management can help determine additional changes or alternative metrics that are required.

Implementation Steps
After agreement on dashboard design, the implementation process may begin. Effective dashboards require live data feeds and, hence, the data integration process may be complex because of multiple data sources. Here is a list of the steps involved in implementation.

• Select final dashboard metrics
• Identify data sources
• Assess feasibility
• Assign metric accountability
• Develop action plan
• Create timeline
• Populate initial metrics
• Establish internal and external benchmarks
• Determine targets
• Determine monitoring, interpretation, feedback procedures and guidelines
• Develop corrective action review process

Best practice allows for online dashboards that may be customized to a users needs. For example, the matrix tool described above might be provided online and the user could select from these metrics those they were interested in and build up there own dashboard. In addition, each user will want the ability to drill down to a level in the organization that is relevant to their position (i.e. a district manager wants to see his district data).

In conclusion, the dashboard design process is detailed and requires thorough research. In addition, data integration and online application development are critical. However, the benefits of an effective dashboard far outweigh the costs in allowing management the critical measures necessary to guide the organization toward success.

Falconeris Marimon Caneda
Socio Director
TTS Consulting

viernes, 17 de septiembre de 2010

The Agile Data Warehouse: Keeping Users Happy


Though they share a single word, agile data warehousing (DW) is nothing like agile software development.

Agile programming disciplines tend to champion a code-first, document-later ethic. Some agile approaches even eschew traditional documentation altogether. Agile programming techniques tend to place an emphasis on frequent testing: at least one agile discipline, test-driven development (TDD), explicitly prescribes a test-first approach.

In all of their variants, agile approaches emphasize the importance of frequent (and typically interactive) involvement with line-of-business customers. It isn't unusual for agile teams to solicit feedback from customers on a periodic (daily, weekly, or bi-weekly) basis. This lets them incorporate new features as customers demand them -- or change features based on feedback from users.

There are a number of reasons why a straight-up agile approach doesn't translate very well into the data warehousing world, experts say.

There's the important paradigmatic distinction between programming -- with its procedural (or line-by-line) orientation -- and data management (DM), which typically lives and thinks in a set-based world.

There are practical logistical concerns, too. "You have to look at it kind of differently, because it can take you longer to write a test case than it takes us to generate the code for you. Suddenly, you're in a different paradigm.

When you're building warehouses in an agile fashion, you're bringing together the concepts of software development and data, and a lot of the agile software techniques don't flow across to the data world."

A lot of the agile buzz at last month's TDWI World Conference in San Diego concerned agile business intelligence (BI), which, Whitehead respectfully suggests, isn't at all the same thing as agile data warehousing.

"When people talk agile in the data world, they generally talk agile BI. They generally talk about the reports, that sort of layer becoming agile. That's a no-brainer. If it's a distinct point where you have customer interaction, of course you should put something in front of them. It isn't quite so easy with a data warehouse," he argues.

All the same, Whitehead describes himself as a proponent of agile data warehousing, particularly inasmuch as "agility" connotes the acceleration or automation of tedious, onerous, time-consuming, or otherwise costly tasks.

Agility is, of course, synonymous with nimbleness, deftness -- that is, speed.

Finally, The essence of agile: "If you're a data guy, you need to make sure that you are doing whatever you can to deliver quickly and deliver value and make changes so that your stuff is relevant, If you can't do that, people are going to fill that vacuum."

viernes, 27 de agosto de 2010

Stage or Not to Stage in Data Warehouse


The back room area of the data warehouse has frequently been called the staging area. Staging in this context means writing to disk and, at a minimum, I recommend staging data at the four major checkpoints of the ETL data flow. But, the main cuestion in this note is: When i need to design stage area in my data warehouse project?


To Stage or Not to Stage

The decision to store data in a physical staging area versus processing it in memory is ultimately the choice of the ETL architect. The ability to develop efficient ETL processes is partly dependent on being able to determine the right balance between physical input and output (I/O) and in-memory processing.

The challenge of achieving this delicate balance between writing data to staging tables and keeping it in memory during the ETL process is a task that must be reckoned with in order to create optimal processes. The issue with determining whether to stage your data or not depends on two conflicting objectives:


- Getting the data from the originating source to the ultimate target as
fast as possible

- Having the ability to recover from failure without restarting from the beginning of the process


The decision to stage data varies depending on your environment and business requirements. If you plan to do all of your ETL data processing in memory, keep in mind that every data warehouse, regardless of its architecture or environment, includes a staging area in some form or another.Consider the following reasons for staging data before it is loaded into the data warehouse:

- Recoverability. In most enterprise environments, it’s a good practice to stage the data as soon as it has been extracted fromthe source system and then again immediately after each of the major transformation steps, assuming that for a particular table the transformation steps are significant. These staging tables (in a database or file system) serve as recovery points. By implementing these tables, the process won’t have to intrude on the source system again if the transformations fail. Also, the process won’t have to transform the data again if the load process fails. When staging data purely for recovery purposes, the data should be stored in a sequential file on the file system rather than in a database. Staging for recoverability is especially important when extracting from operational systems that overwrite their own data.

- Backup. Quite often, massive volume prevents the data warehouse from being reliably backed up at the database level.We’ve witnessed catastrophes that might have been avoided if only the load files were saved, compressed, and archived. If your staging tables are on the file system, they can easily be compressed into a very small footprint and saved on your network. Then if you ever need to reload the data warehouse, you can simply uncompress the load files and reload them.

- Auditing. Many times the data lineage between the source and target is lost in the ETL code. When it comes time to audit the ETL process, having staged data makes auditing between different portions of the ETL processes much more straightforward because auditors (or programmers) can simply compare the original input file with the logical transformation rules against the output file. This staged data is especially useful when the source system overwrites its history. When questions about the integrity of the information in the data warehouse surface days or even weeks after an event has occurred, revealing the staged extract data from the period of time in question can restore the trustworthiness of the data warehouse.


Once you’ve decided to stage at least some of the data, you must settle on the appropriate architecture of your staging area. As is the case with any other database, if the data-staging area is not planned carefully, it will fail. Designing the data-staging area properly is more important than designing the usual applications because of the sheer volume the data-staging area accumulates (sometimes larger than the data warehouse itself).

jueves, 19 de agosto de 2010

What's Essential -- And What's Not -- In Big Data Analytics (Columnar Data Base?)


Far from arguing over the benefits (or drawbacks) of a column-based architecture, shops would be better advised to focus on other, potentially more important issues. Row- or column-based engines marketed by Aster Data, Dataupia, Greenplum Software Inc. (now an EMC Corp. property), Hewlett-Packard Co. (HP), InfoBright, Kognitio, Netezza, ParAccel, Sybase Inc. (now an SAP AG property), Teradata, Vertica, and other vendors (to say nothing of the specialty warehouse configurations marketed by IBM, Microsoft, and Oracle) are by definition architected for Big Analytics.

Analytic database vendors today compete on the basis of several options -- capabilities such as in-database analytics, support for non-traditional (typically non-SQL) query types, sophisticated workload management, and connectivity flexibility.

Every vendor has an option-laden sales pitch, of course -- but few (if any) stories are exactly the same. In-database analytics is particularly hot, according to Eckerson. All analytic database vendors say they support it (to a degree), but some -- such as Aster Data, Greenplum, and (more recently) Netezza, Teradata, and Vertica -- seem to support it "more" flexibly than others.

"With in-database analytics, scoring can execute automatically as new records enter the database rather than in a clumsy two-step process that involves exporting new records to another server and importing and inserting the scores into the appropriate records," he explains.

The twist comes by virtue of (growing) support for non-SQL analytic queries, chiefly in the form of the (increasingly ubiquitous) MapReduce algorithm. Aster Data and Greenplum have supported in-database MapReduce for two years; more recently, both Netezza and Teradata, along with IBM, have announced MapReduce moves. Last month, open source software (OSS) data integration (DI) player Talend announced support for Hadoop (an OSS implementation of MapReduce) in its enterprise DI product. Talend's MapReduce implementation can theoretically support in-database crunching in conjunction with Hadoop-compliant databases.

"[T]echniques like MapReduce make it possible for business analysts, rather than IT professionals, to custom-code database functions that run in a parallel environment," he writes. As implemented by Aster Data and Greenplum, for example, in-database MapReduce permits analysts or developers to write reusable functions in many languages (including the Big Five of Python, Java, C, C++, and Perl) and invoke them by means of SQL calls.

Such flexibility is a harbinger of things to come, according to Eckerson. "[A]s analytical tasks increase in complexity, developers will need to apply the appropriate tool for each task," he notes. "No longer will SQL be the only hammer in a developer's arsenal. With embedded functions, new analytical databases will accelerate the development and deployment of complex analytics against big data."