Showing posts with label quantitative usability. Show all posts
Showing posts with label quantitative usability. Show all posts

Wednesday, October 13, 2010

20101013 - Hu, ...Quantitative Usability Requirements Spec & Usability Evaluation..

A full life-cycle methodology for structured use-centered quantitative usability requirements specification and usability evaluation of websites

by Hu, Guoqiang, Ph.D., Auburn University, 2009 , 201 pages; AAT 3386203



My Interest:

1) QUEST – Hu's usability evaluation method.

2) Expert usability review.

3) User usability testing.

4) SUS – System Usability Scale.

5) Usability metrics (usability criteria) of QUEST.


Action:

To search for his dissertation or journal article. Want to read more.



Background


World Wide Web has gained its dominant status in the cyber information and services delivery world in recent years. But how to specify website usability requirements and how to evaluate and improve website usability according to its usability requirements specification are still big issues to all the stakeholders.


Research Goal


To help solve this problem, we propose a website usability requirements specification and usability evaluation methodology that features a structured use-centered quantitative full life-cycle method.


Methodology


A validation experiment has been designed and conducted to prove the validity of the proposed methodology, QUEST (Quantitative Usability Equations SeT). Its principle is to prove that QUEST has stronger website usability evaluation capability than the most typical existing usability evaluation methods. Apparently, if QUEST's website usability evaluation capability is established, then its usability metrics can be used to quantitatively specify upfront user usability requirements for websites.


In the validation experiment, 7 usability experts and 20 student subjects were recruited to perform 4 tasks on 2 open source calendar websites, WebCalendar 1.0.5 and VCalendar 1.5.3.1; 4 sets of usability data had been collected, which were corresponding to the following 4 usability evaluation methods respectively: expert usability review, traditional user usability testing, SUS (System Usability Scale), and QUEST.


Comments: He should also compare against other popular usability evaluation methods such as heuristic evaluation, and cognitive walkthrough. He should also compare against other more popular Usability Questionnaires, such as QUIS and SUMI. Noted that a good method has been compared, i.e. user usability testing. Generally, he has not done a FAIR comparison.


Results Discussion


According to the experiment results: both the expert usability review and the traditional user usability testing were inconclusive on which of the 2 target websites had better usability; although SUS rated the overall usability of WebCalendar 1.0.5 at 66.00 and VCalendar 1.5.3.1 at 61.75, it was subjective and vague on usability problems; in contrast, QUEST not only rated the overall usability of WebCalendar 1.0.5 at 56.59 and VCalendar 1.5.3.1 at 35.97, but also revealed where the usability problems were and how severe each usability problem was in a quantitative manner.


Comments: After reading Hu's abstract, an idea came to me. Initially, I have thought of selecting the type/hybrid of UEM based on literature review and comparison. Benchmarking on Hu, I could also do a study to compare various UEM and select the UEM type/hybrid based on the results of study.


Conclusion


In conclusion, it clearly can be stated that QUEST has stronger website usability evaluation capability than all other 3 most typical existing usability evaluation methods. So, the proposed methodology has been validated by the experiment results.



Note: No Preview nor Full-Text dissertation is available for download.

20101013 - Yen, ...Usability Evaluation: Methods, Models, Measures

Health information technology usability evaluation: methods, models, and measures

by Yen, Po-Yin, Ph.D., Columbia University, 2010 , 160 pages; AAT 3420882



My Interest:

1) Usability evaluation studies.

2) Usability evaluation scale (Health-ITUES).

3) Usability evaluation model (Health-ITUEM).

4) Validity.


Action:

To read the Dissertation in the future.



Introduction


Health information technology (IT) can offer important benefits to health care; however, technology-related factors are a major obstacle to health IT adoption.


Research Goal


Toward the goal of achieving a greater understanding of health IT usability and its measurement, the dissertation comprised three major analyses:

1) a methodological review of health IT usability evaluation studies to identify problems in existing studies;


2) exploratory factor analysis of the Health Information Technology Usability Evaluation Scale (Health-ITUES) which was developed as part of the dissertation research along with the underlying Health Information Technology Usability Evaluation Model (Health-ITUEM); and


3) confirmatory factor analysis and structural equation modeling to examine the construct validity and predictive validity of Health-ITUES.


Methodology


The health IT system that served as the focus of the analysis was a web-based communication system that supported nurse staffing and scheduling. The sample comprised 553 staff nurses in two healthcare organizations.


In the usability methodological review, we identified problems in existing studies including lack of theoretical framework/model, inconsistent usability definition and evaluation methods, and lack of power analysis for sample size calculation.


Results Discussion


The exploratory factor analysis resulted in a 20-item Health-ITUES comprising four factors that demonstrated strong internal consistency reliability:

quality of work life (QWL), 3 items, α=.94;

perceived usefulness (PU), 9 items, α=.94;

perceived ease of use (PEU), 5 items, α=.95;

user control (UC), 3 items, α=.81.


The confirmatory factor analysis showed that a general usability factor accounted for 78.1%, 93.4%, 51.0% and 39.9% of the explained variance in QWL, PU, PEU, and UC respectively. The structural equation modeling supported the predictive validity of Health-ITUES, explaining 64% of the variance in intention for system use.


Contribution


The results of the dissertation contribute to enhancing the methodological breadth and rigor of health IT usability evaluation studies.



Chapter 1. Introduction 1

Background 2

Definitions of Usability 3

Aspects of Usability 7

Definition and Scope 8

Health IT Usability, Acceptance, and Adoption 10

Health IT Usability Specification and Evaluation 11

Problem Statement 12

Purpose 13

Study Aims and Research Questions 13

Significance of the Study 15


Chapter 2. Methodological Review of Health Information Technology Usability

Specification and Evaluation Studies 16

Background 16

Usability Model 17

System Development Life Cycle 17

An Integrated Usability Specification and Evaluation Framework 18

Methods 23

Search Strategy 23

Inclusion/Exclusion Criteria: 23

Data Extraction and Management , 24

Results 27

Types of Health IT Evaluated 32

Summary of studies categorized in each stage 33

Stage 1: Specify Needs and Setting 35

Stage 2: System Component Development 37

Stage 3: Combination Components 38

Stage 4: Integrate Health IT into the Real Environment 40

Stage 5: Routine Use 41

Study Design and Data Analysis in Stages 4 and 5 41

Discussion 45

Methodological Problems in Existing Studies 45

Objective versus Subjective Measures 47

Environmental Factor Not Evaluated in the Early Stages 49

Inadequate Measure 49

Limitations 50

Conclusion 50


Chapter 3. Health IT Usability Evaluation Model and Scale Development 65

Background 65

Health IT Usability Evaluation Model 70

Definition of Concepts 74

Health IT Usability Evaluation Scale Development 76

The web-based communication system 76

Item selection, creation, and modification 78

Health IT Usability Evaluation Scale Psychometric Evaluation 82

Research Questions 82

Methods 82

Results 84

Discussion 92

Limitations 94

Conclusion 95


Chapter 4. Health-IT Usability Evaluation Scale Confirmatory Analyses 96

Background 96

Research Questions 96

Methods 97

Setting and Sample 97

Sample size 97

Data collection procedures 98

Data Analysis 98

Results 99

Descriptive analysis 99

Power Analysis 100

Construct validity 101

Predictive Validity 104

Discussion 106

Construct and Predictive Validity 106

Methodological issues in existing model testing studies 106

Limitations 107

Conclusions 108

Thursday, September 23, 2010

20100923 - Rihal, ..Objective & Subjective Usability

Relationship between certain objective performance measures and subjective usability

by Rihal, Saravjit Singh, Ph.D., Texas A&M University, 2001 , 190 pages; AAT 3011790



My Interest:

(1) Selecton of 6 objective usability characteristics.

(2) Selection of 1 subjective usability characteristic, i.e. Satisfaction.

(3) Pair-wise comparison.

(4) Shortlisting of objective usability characteristics from 6 to 4.


Action: to read the Dissertation in future.


Motivation


The nature of the relationship between objective and subjective usability measures is thought to be a positive one, though little else is known. It is this relationship that is to be investigated in this study.

Research Objective


The objective is to determine a "figure of merit" (FOM) incorporating the most important objective usability characteristics that can successfully predict the satisfaction characteristic.


Methodology


The usability literature was surveyed to determine the most commonly cited objective and subjective usability characteristics and measures.


Six objective usability characteristics were determined to be important: Learnability, Relearnability, Guessability, Flexibility, Effectiveness, and Efficiency. These objective characteristics coupled with Satisfaction, were presented to experts for pair-wise comparison to determine the perceived importance of each characteristic to usability as a whole.


Analysis of Data


A rank order of the characteristics was obtained and a Kendall's Coefficient of Concordance (Kendall's W) was calculated to determine the level of agreement between the rankings. Little agreement was found among the rankings. One by one, the characteristics were dropped from consideration and the Kendall's W was recalculated to find the combination of characteristics with the highest amount of agreement among the experts. At the end of this process only four objective characteristics remained: Effectiveness, Efficiency, Guessability and Learnability.


Methodology


A web-based usability evaluation study was then performed during which measures of the four characteristics were automatically taken and collected, and compared against the satisfaction measure obtained from the participants of the study.


Results show that the objective characteristics are independent of each other and could predict the satisfaction characteristic.


Comments: In other words, when the objective usability characteristics are high, satisfaction (subjective usability characteristic) is also high.


Conclusion/Discussion

Based on the findings of this study, two proposals are made for a relative Figure of Merit (FOM).

Both FOMs are constructed using the four objective characteristics as determined by the paired comparison; one involves summation, the other multiplication. Both FOMs allow differentiation between the websites and seemingly achieve the objective of predicting user satisfaction.


Comments: Blurr. Wonders what he is talking about when he writes about FOM.



Comments: Year 2001. Softcopy is scanned version; difficult to copy & paste, difficult to blog about it.



Tuesday, September 21, 2010

20100922 - Ivory, ..Automated Web Interface Evaluation

An empirical foundation for automated Web interface evaluation

by Ivory, Melody Yvette, Ph.D., University of California, Berkeley, 2001 , 466 pages; AAT 3044509


My Interest:

(1) Automated web evaluation methodology & tools.

(2) Synthesis of usability & performance evaluation.


Action:

Low priority to read the Dissertation in future.



This dissertation explores the development of an automated Web evaluation methodology and tools. It presents an extensive survey of usability evaluation methods for Web and graphical interfaces and shows that automated evaluation is greatly underexplored, especially in the Web domain.


This dissertation presents a new methodology for HCI: a synthesis of usability and performance evaluation techniques, which together build an empirical foundation for automated interface evaluation.


The general approach involves:

(1) identifying an exhaustive set of quantative interface measures;

(2) computing measures for a large sample of rated interfaces;

(3) deriving statistical models from the measures and ratings;

(4) using the models to predict ratings for new interfaces; and

(5) validating model predictions.


Methodology – Statistical Models


This dissertation presents a specific instantiation for evaluating information-centric Web sites. The methodology entails computing 157 highly-accurate, quantitative page-level and site-level measures. The measures assess many aspects of Web interfaces, including the amount of text on a page, color usage, and consistency. These measures along with expert ratings from Internet professionals are used to derive statistical models of highly-rated Web interfaces. The models are then used in the automated analysis of Web interfaces.


Quantitative Measures of websites


This dissertation presents analysis of quantitative measures for over 5300 Web pages and 330 sites. It describes several statistical models for distinguishing good, average, and poor pages with 93%-96% accuracy and for distinguishing sites with 68%-88% accuracy.


Comments: My impression is that this Dissertation may be uninteresting to read, lots of stats talk. I don't like the Abstract..sigh.


Note: Softcopy Dissertation is scanned version; hence, could not copy and paste; could not blog easily about it.

Friday, September 25, 2009

Sep 25 - Nielsen, Risks of Quantitative Studies (Alertbox)

Summary: Number fetishism leads usability studies astray by focusing on statistical analyses that are often false, biased, misleading, or overly narrow. Better to emphasize insights and qualitative research.

Risks of Quantitative Studies

There are two main types of user research: quantitative (statistics) and qualitative (insights).
The key benefit of quantitative studies is simple: they boil a complex situation down to a single number that's easy to grasp and discuss. I exploit this communicative clarity myself, for example, in reporting that using websites is 206% more difficult for users with disabilities and 122% more difficult for senior citizens than for mainstream users.

Beware Number Fetishism

When I read reports from other people's research, I usually find that their qualitative study results are more credible and trustworthy than their quantitative results. It's a dangerous mistake to believe that statistical research is somehow more scientific or credible than insight-based observational research. In fact, most statistical research is less credible than qualitative studies.

User interfaces and usability are highly contextual, and their effectiveness depends on a broad understanding of human behavior.

Fixating on numbers rather than qualitative insights has driven many usability studies astray. As the following points illustrate, quantitative approaches are inherently risky in a host of ways.

Random Results

Researchers often perform statistical analysis to determine whether numeric results are "statistically significant." By convention, they deem an outcome significant if there is less than 5% probability that it could have occurred randomly rather than signifying a true phenomenon.
This sounds reasonable, but it implies that one out of twenty "significant" results might be random if researchers rely purely on quantitative methods.

Luckily, most good researchers -- especiaally those in the user-interface field -- use more than a simple quantitative analysis. Thus, they typically have insights beyond simple statistics when they publish a paper, which drives down, but doesn't eliminate, bogus findings.

There's a reverse phenomenon as well: Sometimes a true finding is statistically insignificant because of the experiment's design. Perhaps the study didn't include enough participants to observe a major -- but rare -- finding in sufficient numbers. It would therefore be wrong to dismiss issues as irrelevant just because they don't show up in quantitative study results.

Pulling Correlations Out of a Hat

If you measure enough variables, you will inevitably discover that some seem to correlate. Run all your stats through the software and a few "significant" correlations will surely pop out. (Remember: one out of twenty analyses are "significant," even if there is no underlying true phenomenon.)

Studies that measure seven metrics will generate twenty-one possible correlations between the variables. Thus, on average, such studies will have one bogus correlation that the statistics program deems "significant," even if the issues being measured have no real connection.
In my Web Usability 2004 project, we collected metrics on fifty-three different aspects of user behavior on websites. There are thus 1,378 possible correlations that I could throw into the hopper. Even if we didn't discover anything at all in the study, about sixty-nine correlations would emerge as "statistically significant."

Overlooking Covariants

Even when a correlation represents a true phenomenon, it can be misleading if the real action concerns a third variable that is related to the two you're studying.

For example, studies show that intelligence declines by birth order. In other words, a person who was a first-born child will on average have a higher IQ than someone who was born second. Third-, fourth-, fifth-born children and so on have progressively lower average IQs. This data seems to present a clear warning to prospective parents: Don't have too many kids, or they'll come out increasingly stupid. Not so.
There's a hidden third variable at play: smarter parents tend to have fewer children. When you want to measure the average IQ of first-born children, you sample the offspring of all parents, regardless of how many kids they have. But when you measure the average IQ of fifth-born children, you're obviously sampling only the offspring of parents who have five or more kids. There will thus be a bigger percentage of low-IQ children in the latter sample, giving us the true -- but misleading -- conclusion that fifth-born children have lower average IQs than first-born children. Any given couple can have as many children as they want, and their younger children are unlikely to be significantly less intelligent than their older ones. When you measure intelligence based on a random sample from the available pool of children, however, you're ignoring the parents, who are the true cause of the observed data.

(Update added 2007: The newest research suggests that there may actually be a tiny advantage in IQ for first-born children after correcting for family size and the parents' economic and educational status. But the point remains that you have to correct for these covariants, and when you do so, the IQ difference is much less than plain averages may lead you to believe.)

As a Web example, you might observe that longer link texts are positively correlated with user success. This doesn't mean that you should write long links. Website designers are the hidden covariant here: clueless designers tend to use short text links like "more," "click here," and made-up words.

Over-Simplified Analysis

To get good statistics, you must tightly control the experimental conditions -- often so tightly that the findings don't generalize to real problems in the real world.
This is a common problem for university research, where the test subjects tend to be undergraduate students rather than mainstream users. Also, instead of testing real websites with their myriad contextual complexities, many academic studies test scaled-back designs with a small page count and simplified content.

For example, it's easy to run a study that shows breadcrumbs are useless: just give users directed tasks that require them to go in a straight line to the desired destination and stop there. Such users will (rightly) ignore any breadcrumb trail. Breadcrumbs are still recommended for many sites, of course. Not only are they lightweight, and thus unlikely to interfere with direct-movement users, but they're helpful to users who arrive deep within a site via search engines and direct links. Breadcrumbs give these users context and help users who are doing comparisons by offering direct access to higher levels of the information architecture.

Usability-in-the-large is often neglected by narrow research that doesn't consider, for example, revisitation behavior, search engine visibility, and multi-user decision-making.

Distorted Measurements

It's easy to prejudice a usability study by helping the users at the wrong time or by using the wrong tasks. In fact, you can prove virtually anything you want if you design the study accordingly. This is often a factor behind "sponsored" studies that purport to show that one vendor's products are easier to use than a competitor's products.

Even if the experimenters aren't fraudulent, it's easy to get hoodwinked by methodological weaknesses, such as directing the users' attention to specific details on the screen.
The very fact that you're asking about some design elements rather than others makes users notice them more and thus changes their behavior.

Many Web advertising studies are misleading, possibly because most such studies come from advertising agencies.
The most common distortion is the novelty effect: whenever a new advertising format is introduced, it's always accompanied by a study showing that the new type of ad generates more user clicks. Sure, that's because the new format enjoys a temporary advantage: it gathers user attention simply because it's new and users have yet to train themselves to ignore it.
The study might be genuine as far as it goes, but it says nothing about the new advertising format's long-term advantages once the novelty effect wears off.

Publication Bias

Editors follow the "man bites dog" principle to highlight new and interesting stories. While understandable, this preference for new and different findings imposes a significant bias in the results that get exposure.

Usability is a very stable field. User behavior is pretty much the same year after year. I keep finding the same results in study after study, as do many others. Every now and then, a bogus result emerges and publication bias ensures that it gets much more attention than it deserves.

Consider the question of Web page download time. Everyone knows that faster is better. Interaction design theory has documented the importance of response times since 1968, and this importance has been seen empirically in countless Web studies since 1995. E-commerce sites that speed up response times sell more. The day your server is slow, you lose traffic. (This happened to me recently: on January 14, Tog got "slashdotted"; because we share a server, my site lost 10% of its normal pageviews for a Wednesday when AskTog's increased traffic slowed useit.com down.)
If twenty people study download times, nineteen will conclude that faster is better. But again: one of every twenty statistical analyses will give the wrong result, and this one study might be widely discussed simply because it's new. The nineteen correct studies, in contrast, might easily escape mention.

Judging Bizarre Results

Bizarre results are sometimes supported by seemingly convincing numbers. You can use the issues I've raised here as a sanity check: Did the study pull correlations out of a hat? Was it biased or overly narrow? Was it promoted purely because it's different? Or was it just a fluke?
Typically, you'll discover that deviant findings should be ignored.
The broad concepts of human behavior in interactive systems are stable and easy to understand. The exceptions usually turn out to be exactly that: exceptions.

In 1989, for example, I published a paper on discount usability engineering, stating that small, fast user studies are superior to larger studies, and that testing with about five users is typically sufficient.
This was quite contrary to the prevailing wisdom at the time, which was dominated by big-budget testing. During the fifteen years since my original claim, several other researchers reached similar conclusions, and we developed a mathematical model to substantiate the theory behind my empirical observation. Today, almost everyone who does user testing has concluded that they learn most of what they'll ever learn with about five users.

But four or five studies constitute a trend, which much enhances the finding's credibility as a general phenomenon.

Quantitative Studies: Intrinsic Risks

All the reasons I've listed for quantitative studies being misleading indicate bad research; it's possible to do good quantitative research and derive valid insights from measurements. But doing so is expensive and difficult.
Quantitative studies must be done exactly right in every detail or the numbers will be deceptive. There are so many pitfalls that you're likely to land in one of them and get into trouble.

If you rely on numbers without insights, you don't have backup when things go wrong. You'll stumble down the wrong path, because that's where the numbers will lead.

Qualitative studies are less brittle and thus less likely to break under the strain of a few methodological weaknesses. Even if your study isn't perfect in every last detail, you'll still get mostly good results from a qualitative method that relies on understanding users and their observed behavior.
Yes, experts get better results than beginners from qualitative studies.

But for quantitative studies, only the best experts get any valid results at all, and only then if they're extremely careful.

Source:
Jakob Nielsen's Alertbox, March 1, 2004:
Risks of Quantitative Studies
http://www.useit.com/alertbox/20040301.html
Risks of Quantitative Studies (Jakob Nielsen's Alertbox)

Tuesday, September 22, 2009

Sep 22 - Putting A/B Testing in Its Place (Alertbox)

Putting A/B Testing in Its Place

Summary: Measuring the live impact of design changes on key business metrics is valuable, but often creates a focus on short-term improvements. This near-term view neglects bigger issues that only qualitative studies can find.

Introduction
In A/B testing, you unleash two different versions of a design on the world and see which performs the best. For decades, this has been a classic method in direct mail, where companies often split their mailing lists and send out different versions of a mailing to different recipients. A/B testing is also becoming popular on the Web, where it's easy to make your site show different page versions to different visitors.
Sometimes, A and B are directly competing designs and each version is served to half the users. Other times, A is the current design and serves as the control condition that most users see. In this scenario, B, which might be more daring or experimental, is served only to a small percentage of users until it has proven itself.

Benefits

Compared with other methods, A/B testing has four huge benefits:

1. It measures the actual behavior of your customers under real-world conditions. You can confidently conclude that if version B sells more than version A, then version B is the design you should show all users in the future.

2. It can measure very small performance differences with high statistical significance because you can throw boatloads of traffic at each design. The sidebar shows how you can measure a 1% difference in sales between two designs.

3. It can resolve trade-offs between conflicting guidelines or qualitative usability findings by determining which one carries the most weight under the circumstances.
For example, if an e-commerce site prominently asks users to enter a discount coupon, user testing shows that people will complain bitterly if they don't have a coupon because they don't want to pay more than other customers. At the same time, coupons are a good marketing tool, and usability for coupon holders is obviously diminished if there's no easy way to enter the code.
When e-commerce sites have tried A/B testing with and without coupon entry fields, overall sales typically increased by 20-50% when users were not prompted for a coupon on the primary purchase and checkout path. Thus, the general guideline is to avoid prominent coupon fields.

4. It's cheap: once you've created the two design alternatives (or the one innovation to test against your current design), you simply put both of them on the server and employ a tiny bit of software to randomly serve each new user one version or the other. Also, you typically need to cookie users so that they'll see the same version on subsequent visits instead of suffering fluctuating pages, but that's also easy to implement. There's no need for expensive usability specialists to monitor each user's behavior or analyze complicated interaction design questions. You just wait until you've collected enough statistics, then go with the design that has the best numbers.

Limitations

With these clear benefits, why don't we use A/B testing for all projects? Because the downsides usually outweigh the upsides.

First, A/B testing can only be used for projects that have one clear, all-important goal, that's to say a single KPI (key performance indicator). Furthermore, this goal must be measurable by computer, by counting simple user actions. Examples of measurable actions include:
Sales for an e-commerce site.
Users subscribing to an email newsletter.
Users opening an online banking account.
Users downloading a white paper, asking for a salesperson to call, or otherwise explicitly moving ahead in the sales pipeline.

For many sites, the ultimate goals are not measurable through user actions on the server. Goals like improving brand reputation or supporting the company's public relations efforts can't be measured by whether users click a specific button.
Similarly, while you can easily measure how many users sign up for your email newsletter, you can't assess the equally important issue of how they read your newsletter content without observing subscribers as they open the messages.

A second downside of A/B testing is that it only works for fully implemented designs. It's cheap to test a design once it's up and running, but we all know that implementation can take a long time.

In contrast, paper prototyping lets you try out several different ideas in a single day. Of course, prototype tests give you only qualitative data, but they typically help you reject truly bad ideas quickly and focus your efforts on polishing the good ones.

Short-Term Focus

A/B testing's driving force is the number being measured as the test's outcome. Usually, this is an immediate user action, such as buying something. In theory, there's no reason why the metric couldn't be a long-term outcome, such as total customer value over a five-year period. In practice, however, such long-term tracking rarely occurs. Nobody has the patience to wait years before they know whether A or B is the way to go.
Basing your decisions on short-term numbers, however, can lead you astray. A common example: Should you add a promotion to your homepage or product pages?

No Behavioral Insights

The biggest problem with A/B testing is that you don't know why you get the measured results. You're not observing the users or listening in on their thoughts. All you know is that, statistically, more people performed a certain action with design A than with design B. Sure, this supports the launch of design A, but it doesn't help you move ahead with other design decisions.

Say, for example, that you tested two sizes of Buy buttons and discovered that the big button generated 1% more sales than the small button. Does that mean that you would sell even more with an even bigger button? Or maybe an intermediate button size would increase sales by 2%. You don't know, and to find out you have no choice but to try again with another collection of buttons.

Of course, you also have no idea whether other changes might bring even bigger improvements, such as changing the button's color or the wording on its label. Or maybe changing the button's page position or its label's font size, rather than changing the button’s size, would create the same or better results.

Worst of all, A/B testing provides data only on the element you're testing. It's not an open-ended method like user testing, where users often reveal stumbling blocks you never would have expected.

Combining Methods

A/B testing has more problems than benefits. You should not make it the first method you choose for improving your site's conversion rates. And it should certainly never be the only method used on a project.
Qualitative observation of user behavior is faster and generates deeper insights. Also, qualitative research is less subject to the many errors and pitfalls that plague quantitative research.

A/B testing does have its own advantages, however, and provides a great supplement to qualitative studies. Once your company's commitment to usability has grown to a level where you're regularly conducting many forms of user research, A/B testing definitely has its place in the toolbox.

Source:
Jakob Nielsen's Alertbox, August 15, 2005:
Putting A/B Testing in Its Place
http://www.useit.com/alertbox/20050815.html
Putting A/B Testing in Its Place (Jakob Nielsen's Alertbox)