Developing heuristic evaluation methods for large screen information exhibits based on critical parameters by Somervell, Jacob, Ph.D., Virginia Polytechnic Institute and State University, 2004 , 235 pages; AAT 3136384 My Interest: 1) Set of heuristics for LSIE. 2) Process of creating the set of heuristics. 3) Method of testing/validating the set of heuristics. 4) Heuristic evaluation methods. Action: To read the Dissertation in detail in future; this may be a benchmark for my research. Motivation Evaluation is the key to effective interface design. It becomes even more important when the interfaces are for cutting edge technology, in application areas that are new and with little prior design knowledge. Knowing how to evaluate new interfaces can decrease development effort and increase the returns on resources spent on formative evaluation. Problem The problem is that there are few, if any, readily available evaluation tools for these new interfaces. Research Goal This work focuses on the creation and testing of a new set of heuristics that are tailored to the large screen information exhibit (LSIE) system class. Set of Heuristics This new set is created through a structured process that relies upon critical parameters associated with the notification systems design space. By inspecting example systems, performing claims analysis, categorizing claims, extracting design knowledge, and finally synthesizing heuristics; we have created a usable set of heuristics that is better equipped for supporting formative evaluation. Contribution of Knowledge Contributions of this work include: a structured heuristic creation process based on critical parameters, a new set of heuristics tailored to the LSIE system class, reusable design knowledge in the form of claims and high level design issues, and a new usability evaluation method comparison test. These contributions result from the creation of the heuristics and two studies that illustrate the usability and utility of the new heuristics. 2.2 Evaluation of Large Screen Information Exhibits . . . . . . . . . . . . . . . . . . 16 2.2.1 Analytical Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 2.2.2 Heuristic Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 2.2.3 Comparing UEMs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 2.2.4 Comparing Heuristics . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 3 Background and Motivation 23 3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 3.2 Assessing Evaluation Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 3.3 Motivation from Prior Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 3.4 Experiment Description . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 3.4.1 System Descriptions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 3.4.2 Methodology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 26 3.5 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 3.5.1 Drawbacks to Surveys . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29 3.5.2 Strength of Claims Analysis . . . . . . . . . . . . . . . . . . . . . . . . . 30 3.6 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 4 Heuristic Creation 32 4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 4.2 Motivation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34 4.3 Processes Involved . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34 4.4 Selecting Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 4.4.1 Are these LSIEs? . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 36 4.4.2 Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 4.5 Analyzing Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 4.5.1 Claims Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 4.5.2 System Claims . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 4.5.3 Validating Claims . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 4.6 Categorizing Claims . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 43 4.6.1 Classifying Claims Using the IRC Framework . . . . . . . . . . . . . . . . 43 4.6.2 Assessing Goal Impact . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 4.6.3 Categorization Through Scenario Based Design . . . . . . . . . . . . . . . 45 4.7 Synthesis Into Heuristics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 49 4.7.1 Visualizing the Problem Tree . . . . . . . . . . . . . . . . . . . . . . . . . 49 4.7.2 Identifying Issues . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 4.7.3 Issues to Heuristics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 53 4.7.4 Heuristics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54 4.8 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55 4.9 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 56 5 Heuristic Comparison Experiment 58 5.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 5.2 Approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 5.2.1 Heuristic Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59 5.2.2 Comparison Technique . . . . . . . . . . . . . . . . . . . . . . . . . . . . 61 5.2.3 Systems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62 5.2.4 Hypotheses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 66 5.2.5 Identifying Problem Sets . . . . . . . . . . . . . . . . . . . . . . . . . . . 66 5.3 Testing Methodology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68 5.3.1 Participants . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 68 5.3.2 Materials . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69 5.3.3 Questionnaire . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69 5.3.4 Measurements Recorded . . . . . . . . . . . . . . . . . . . . . . . . . . . 70 5.4 Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70 5.4.1 Participant Experience . . . . . . . . . . . . . . . . . . . . . . . . . . . . 70 5.4.2 Applicability Scores . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 72 5.4.3 Thoroughness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 74 5.4.4 Validity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 5.4.5 Effectiveness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 77 5.4.6 Reliability – Differences . . . . . . . . . . . . . . . . . . . . . . . . . . . 78 5.4.7 Reliability – Agreement . . . . . . . . . . . . . . . . . . . . . . . . . . . 81 5.4.8 Time Spent . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 5.5 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 5.5.1 Hypotheses Revisited . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 5.5.2 Other Discussion and Implications . . . . . . . . . . . . . . . . . . . . . . 86 5.6 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 87 6 Heuristic Application 88 6.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 88 6.2 Novice HCI Students . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 88 6.2.1 Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89 6.2.2 Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89 6.2.3 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 90 6.2.4 Post-Analysis of Problems . . . . . . . . . . . . . . . . . . . . . . . . . . 92 6.2.5 Evaluator Ability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 92 6.3 Education Domain Experts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93 6.3.1 Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 94 6.3.2 Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 94 6.3.3 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 6.3.4 Post-Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 6.3.5 GAWK Re-Design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96 6.4 HCI Expert Opinions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97 6.4.1 Method . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 97 6.4.2 Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98 6.4.3 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98 6.5 Overall Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98 6.6 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98 Comments: I hope to read and understand how he did it. May be a benchmark for me. |
Showing posts with label Somervell. Show all posts
Showing posts with label Somervell. Show all posts
Thursday, September 23, 2010
20100924 - Somervell, ..Heurics Evaluation Methods for LSIE...
Sunday, November 8, 2009
Week of Nov 2-7: Progress Report
Was in Jakarta from Nov 1 (Sun) till Nov 7 (Sat).
Had achieved a lot of thorough reading of:
Rubin, Jeffrey. Handbook of Usability Testing.
Somervell, Jacob. Developing Heuristic Evaluation Methods for Large Screen Information Exhibits Based on Critical Parameters. [Dissertation, PhD in Computer Science and Applications]
Baker, Kevin F. Heuristic Evaluation of Shared Workspace Groupware based on the Mechanics of Collaboration. [Thesis, M.Sc.]
Good benchmark for me to write my chapters 1 and 2.
Through Rubin's book, I have understood more deeply about who, what, when and how of usability testing.
Had achieved a lot of thorough reading of:
Rubin, Jeffrey. Handbook of Usability Testing.
Somervell, Jacob. Developing Heuristic Evaluation Methods for Large Screen Information Exhibits Based on Critical Parameters. [Dissertation, PhD in Computer Science and Applications]
Baker, Kevin F. Heuristic Evaluation of Shared Workspace Groupware based on the Mechanics of Collaboration. [Thesis, M.Sc.]
Good benchmark for me to write my chapters 1 and 2.
Through Rubin's book, I have understood more deeply about who, what, when and how of usability testing.
Wednesday, November 4, 2009
Nov 5 - Somervell, Heuristic Comparison Experiment (PhD dissertation)
Chapter 5
Heuristic Comparison Experiment
5.1 Introduction
Now that there is a set of heuristics tailored for the large screen information exhibit system class, a comparison of this set to more established types of heuristics can be done. The purpose of this comparison would be to show the utility of this new heuristic set. This comparison needs to be fair, so that determining the effectiveness of the new method will be accurate.
To assess whether the new set of heuristics provides better usability results over existing alternative sets, we conducted a comparison experiment in which each of three sets of heuristics were used to evaluate three separate large screen information exhibits. We then compared the results of each set through several metrics to determine the better evaluation methods for large screen information exhibits.
5.2 Approach
The following sections provide descriptions of the heuristics used, the comparison method, and the systems used in this experiment.
5.2.1 Heuristic Sets
We used three different sets of usability heuristics, each at a different level of specificity for application to large screen information exhibits, ranging from a set completely designed for this particular system class, to a generic set applicable to a wide range of interactive systems.
Nielsen
The least specific set of heuristics was taken from Nielsen and Mack [70]. This set is intended for use on any interactive system, mostly targeted towards desktop applications. Furthermore, this set has been in use since around 1990. It has been tested and criticized for years, but still remains popular with usability practitioners. Again, this set is not tailored for large screen information exhibits in any way and has no relation to the critical parameters for notification systems.
Visibility of system status
Match between system and real world
User control and freedom
Consistency and standards
Error prevention
Recognition rather than recall
Flexibility and efficiency of use
Aesthetic and minimalist design
Help users recognize, diagnose, and recover from errors
Help and documentation
Figure 5.1: Nielsen’s heuristics. General heuristics that apply to most interfaces. Found in [70].
Berry
The second heuristic set used in this comparison test was created for general notification systems by Berry [9]. This set is based on the critical parameters associated with notification systems [62], but only at cursory levels. This set is more closely tied to large screen information exhibits than Nielsen’s method in that large screen information exhibits are a subset of notification systems, but this set is still generic in nature with regards to the specifics surrounding the LSIE system class.
Notifications should be timely
Notifications should be reliable
Notification displays should be consistent (within priority levels)
Information should be clearly understandable by the user
Allow for shortcuts to more information
Indicate status of notification system
Flexibility and efficiency of use
Provide context of notifications
Allow adjustment of notification parameters to fit user goals
Figure 5.2: Berry’s heuristics. Tailored more towards Notification Systems in general. Found in
[9].
Somervell
The final heuristic set is the one created in this work, as reported in Chapter 4. This set is tailored specifically to large screen information exhibits, and thus would be the most specific method of the three when targeting this type of system. It is based on specific levels of the critical parameters associated with the LSIE system class.
5.2.2 Comparison Technique
To determine which of the three sets is better suited for formative evaluation of large screen information exhibits, we use a current set of comparison metrics that rely upon several measures of a method’s ability to uncover usability problems through an evaluation.
The comparison method we are using typically relies on five separate measures to assess the
utility of a given UEM for one’s particular needs, but we will only use a subset in this particular
comparison study.
Hartson et al. report that thoroughness, validity, effectiveness, reliability, and
downstream utility are appropriate measures for comparing evaluation methods [40].
Specifically, our comparison method capitalizes on thoroughness, validity, effectiveness, and reliability, abandoning the downstream utility measure. This choice is used here because longterm studies are required to illustrate downstream utility.
Thoroughness
This measure gives an indication of a method’s ability to uncover a significant percentage of the problems in a given system. Thoroughness consists of a simple calculation of the number of problems uncovered by a single UEM divided by the total number of problems found by all three methods.
thoroughness = (# of problems found by target UEM) / (# of problems found by all methods)
Validity
Validity refers to the ability of a method to uncover the types of problems that real users would experience in day to day use of the system, as opposed to simple or minor problems. Validity is measured as the number of real problems found divided by the total number of real problems identified in the system.
validity = (# of problems found by target UEM) / (# of problems in the system)
The number of real problems in the system refers to the problem set identified through some standard method that is separate from the method being tested.
Effectiveness
Effectiveness combines the previous two metrics into a single assessment of the method. This measure is calculated by multiplying the thoroughness score by the validity score.
effectiveness = thoroughness X validity
Reliability
Reliability is a measure of the consistency of the results of several evaluators using the method. This is also sometimes referred to as inter-rater reliability. This measure is taken more as agreement between the usability problem sets produced by different people using a given method.
This measure is calculated from the differences among all of the evaluators for a specific system as well as by the total number of agreements among the evaluators, thus two measures are used to provide a more robust measurement of the reliability of the heuristic sets:
reliability-d = difference among evaluators for a specific method
reliability-a = average agreement among evaluators for a specific method
For calculating reliability, Hartson et al. recommend using a method from Sears [81] that depends on the ratio of the standard deviation of the numbers of problems found by the average number found [40]. This measure of reliability is overly complicated for current needs, thus a more traditional measure that relies upon actual rater differences is used instead.
5.2.3 Systems
Three systems were used in the comparison study providing a range of applications for which each heuristic would be used in an analytic evaluation. The intent was to provide enough variability in the test systems to tease out differences in the methods.
1 Source Viewer
2 Plasma Poster
3 Notification Collage
Why Source Viewer?
The Source Viewer was chosen as a target system for this study because we wanted an example of a real system that has been in regular use for an extended period. We immediately thought of command and control situations. Potential candidates included local television stations, local air traffic control towers, electrical power companies, and telephone exchange stations. We finally settled on local television command and control after limited responses from the other candidates.
Why Plasma Poster?
We wanted to include the Plasma Poster because it is one of very few LSIE systems that has seen some success in terms of long term usage and acceptance. It has seen over a year of deployment in a large research laboratory, with reports on usage and user feedback reported in [20].
This lengthy deployment and data collection period provides ample evidence for typical usability problems. We can use the published reports as support for our problem sets. Coupled with developer feedback, we can effectively validate the problem set for this system.
Why Notification Collage?
We chose the Notification Collage as the third system for several reasons.
First, we wanted to increase the validity of any results we find. By using more systems, we get a better picture of the “goodness” of the heuristic sets, especially if we get consistent results across all three systems.
Secondly, we wanted to explicitly show that the heuristic set we created in this work actually uncovered the issues that went into that creation process. In other words, since the Notification Collage was one of the systems that led to this heuristic set, using that set on the Notification Collage should uncover most of the issues with that system.
Finally, we wanted to use the Notification Collage out of the original five because we had the most developer feedback on that system, and like the Plasma Poster, it has seen reasonable deployment and use.
5.2.4 Hypotheses
We have three main hypotheses to test in this experiment:
1. Somervell’s set of heuristics has a higher validity score for the Notification Collage.
We believed this was true because the Notification Collage was used in the creation of Somervell’s heuristics, thus those heuristics should identify most or all of the issues in the Notification Collage.
2. More specific heuristics have higher thoroughness, validity, and reliability measures.
We felt this was true because more specific methods are more closely related to the systems in this study. Indeed, from Chapter 3, we discussed how previous work suggests system-class level heuristics would be best. This experiment illustrates this case for heuristic evaluation of large screen information exhibits.
3. Generic methods require more time for evaluators to complete the study.
This seems logical because a more generic heuristic set would require more interpretation and thought, hence we felt that those evaluators who use Nielsen’s set would take longer to complete the system evaluations, providing further impetus for developing system-class UEMs.
5.2.5 Identifying Problem Sets
One problem identified in other UEM comparison studies involves the calculation of specific metrics that rely upon something referred to as the “real” problem set (see [40]). In most cases, this problem set is the union of the problems found by each of the methods in the comparison study. In other words, each UEM is applied in a standard usability evaluation of a system, and the “real” problem set is simply the union of the problems found by each of the methods.
This comparison study also faced the same challenge. Instead of relying on evaluators to produce sets of problems from each method, then using the union of those problem sets as the “real” problem set, analysis and testing was performed on the target systems beforehand and the problem reports from those efforts were used to come up with a standard set of problems for each system.
Source Viewer Problem Set
To determine the problem sets experienced by the users of this system, a field study was conducted. Two interviews with the users of the large screen system, as well as observation were conducted.
Plasma Poster Problem Set
Analytic evaluation augmented with developer feedback and literature review served as the method for determining the real problem set for the Plasma Poster. We employed the same claims analysis technique that we used in the creation process to identify typical usability tradeoffs for the Plasma Poster. After identifying the usability issues, we asked the developers of the system to verify the tradeoffs.
Notification Collage Problem Set
To validate the problem set for the Notification Collage, we contacted the developers of the system and asked them to check each tradeoff as it pertained to the behavior of real users. The developers were given a list of the tradeoffs found in our claims analysis (from Chapter 4) and asked to verify each tradeoff according to their observations of real user behavior.
5.3 Testing Methodology
This experiment involves a 3x3 mixed factors design. We have three levels of heuristics (Nielsen, Berry, and Somervell) and three systems (Source Viewer, Plasma Poster, and Notification Collage).
The heuristics variable is a between-subjects variable because each evaluator sees only one set of heuristics. The system variable is within-subjects because each participant sees all three systems. For example, evaluator 1 saw only Nielsen’s heuristics, but used those to evaluate all three systems.
We used a balanced Latin Square to ensure learning effects from system presentation order would be minimized. Thus, we needed a minimum of 18 participants (6 per heuristic set) to ensure coverage of the systems in the Latin Square balancing.
5.3.1 Participants
As shown in Table 5.1, we needed a minimum of 18 evaluators for this study.
Twenty-one computer science graduate students who had completed a course on Usability Engineering volunteered for participation as inspectors. Six participants were assigned to each heuristic set, to cover each of the order assignments. Three additional students volunteered and they were randomly assigned a presentation order.
These participants all had knowledge of usability evaluation, as well as analytic and empirical methods. Furthermore, each was familiar with heuristic evaluation. Some of the participants were not familiar with the claim structure used in this study, but they were able to understand the tradeoff concept immediately.
Unfortunately, one of the participants failed to complete the experiment. This individual apparently decided the effort required to complete the test was too much, and thus filled out the questionnaire using a set pattern.
This makes the final number of participants 20, with seven for Nielsen’s heuristics, seven for Berry’s heuristics, and six for Somervell’s heuristics.
5.3.2 Materials
Each target system was described in one to three short scenarios, and screen shots were provided to the evaluators. The goal was to provide the evaluators with a sense of the display and its intended usage.
This material is sufficient for the heuristic inspection technique according to Nielsen and Mack [70]. This setup ensured that each of the heuristic sets would be used with the same material, thereby reducing the number of random variables in the execution of this experiment.
A description of the heuristic set to be used was also provided to the evaluators. This description included a listing of the heuristics and accompanying text clarification. This clarification helps a person understand the intent and meaning of a specific heuristic, hopefully aiding in assessment. These descriptions were taken from [70] and [9] for Nielsen and Berry respectively.
Armed with the materials for the experiment, the evaluator then proceeded to rate each of the heuristics using a 7-point Likert scale, based on whether or not they felt that the heuristic applied to a claim describing a design tradeoff in the interface. Thus they are judging whether or not a specific heuristic applies to the claim, and how much so.
Marks of four or higher indicate agreement that the heuristic applies to the claim, otherwise the evaluator is indicating disagreement that the heuristic applies.
5.3.3 Questionnaire
As mentioned earlier, the evaluators in this experiment provided their feedback through a Likert scale, with agreement ratings for each of the heuristics in the set. In addition to this feedback, each evaluator also rated the claim in terms of how much they felt it actually applied to the interface in question.
By indicating their agreement level with the claim to the interface, we get feedback on whether usability experts actually think the claim is appropriate for the interface in question.
After rating each of the heuristics for the claims, we also asked each evaluator to rank the severity that the claim would hold, if the claim were indeed a usability problem in the interface.
5.3.4 Measurements Recorded
The data collected in this experiment consists of each evaluator’s rating of the claim applicability, each heuristic rating for an individual claim, and the evaluator’s assessment of the severity of the usability problem. This data was collected for each of the thirty claims across the three systems.
In addition to the above measures, we also collected data on the evaluator’s experience with usability evaluation, heuristics, and large screen information exhibits. This evaluator information was collected through survey questions before the evaluation was started.
After the evaluators completed the test, they recorded the amount of time they spent on the task. This was a self reported value as each evaluator worked at his/her own pace and in their own location.
5.4 Results
Twenty-one evaluators provided feedback on 33 different claims across three systems. Each evaluator ended up providing either 10 or 12 question responses per claim, depending on the heuristic set used (Nielsen’s set has 10 in it, whereas the others only have 8). This means we have either 330 or 396 answers to consider, per evaluator.
Fortunately, this data was separable into manageable chunks, dealing with applicability, severity, and heuristic ratings; as well as evaluator experience levels and time to complete for each method.
5.4.1 Participant Experience
As for individual evaluator abilities, the average experience level with usability evaluation, across all three systems, was “amateur”. This means that overall, for each heuristic set, we had comparable experience for the evaluators assigned to that set.
5.4.2 Applicability Scores
To indicate whether or not a heuristic set applied to a given claim (or problem), evaluators marked their agreement with the statement “the heuristic applies to the claim”. This agreement rating indicates that a specific heuristic applied to the claim. Each of the heuristics was marked on a 7-point Likert scale by the evaluators, indicating his/her level of agreement with the statement.
Using this applicability measure, the responses were averaged for a single claim across all of the evaluators. Averaging across evaluators allows assessment of the overall “applicability” of the heuristic to the claim.
This applicability score is used to determine whether any of the heuristics applied to the issue described in the claim. If a heuristic received an “agree” rating, average greater than or equal to five, then that heuristic was thought to have applied to the issue in the claim.
Overall Applicability
Considering all 33 claims together (found in all three systems), one-way analysis of variance (ANOVA) indicates significant differences among the three heuristic sets for applicability (F2,855) = 3.0,MSE = 49.7, p < 0.05). Further pair-wise t-tests reveal that Somervell’s set of heuristics had significantly higher applicability ratings over both Berry’s (df = 526, t = 3.32, p < 0.05) and Nielsen’s sets (df = 592, t = 11.56, p < 0.05). In addition, Berry’s heuristics had significantly higher applicability scores over Nielsen’s set (df = 592, t = 5.94, p < 0.05).
5.4.3 Thoroughness
Recall that thoroughness is measured as the number of problems found by a single method, divided by the number uncovered by all of the methods. This requires a breakdown of the total number of claims into the numbers for each system.
Plasma Poster has 14 claims, Notification Collage has eight claims, and the Source Viewer has 11 claims. We look at thoroughness measures for each system. To calculate the thoroughness measures for the data we have collected, we count the number of claims “covered” by the target heuristic set.
Here we are defining covered to mean that at least one of the heuristics in the set had an average agreement rating of at least five. Why five?
On the Likert scale, five indicates somewhat agree. If we require that the average score across all of the evaluators to be greater than or equal to five for a single heuristic, we are only capturing those heuristics that truly apply to the claim in question.
Overall Thoroughness
Across all three heuristic sets, 28 of 33 claims had applicability scores higher than five. Somervell’s heuristics had the highest thoroughness rating of the three heuristic sets with 96% (27 of 28 claims). Berry’s heuristics came next with a thoroughness score of 86% (24 of 28) and Nielsen’s heuristics had a score of 61 (17 of 28).
5.4.4 Validity
Validity measures the UEM’s ability to uncover real usability problems in a system [40].
Here the full set of problems in the system is used as the real problem set (as discussed in earlier sections).
As with thoroughness, the applicability scores determine the validity each heuristic set held for the three systems. As before, we used the cutoff value of five on the Likert scale to indicate applicability of the heuristic to the claim. An average rating of five or higher indicates that the heuristic applied to the claim in question.
Overall Validity
Similar to thoroughness, validity scores were calculated across all three systems. Out of 33 total claims, only 28 showed applicability scores greater than five across all three heuristic sets. Somervell’s heuristics had the highest validity, with 27 of 33 claims yielding applicability scores greater than five, for a validity score of 82%. Berry’s heuristics had the next highest validity with 24 of 33 claims, for a validity score of 73%. Nielsen’s heuristics had the lowest validity score, with 17 of 33 claims for a score of 52%.
5.4.5 Effectiveness
Effectiveness is calculated by multiplying thoroughness by validity. UEMs that have high thoroughness and high validity will have high effectiveness scores. A low score on either of these measures will reduce the effectiveness score.
Overall Effectiveness
Considering the effectiveness scores across all three systems reveals that Somervell’s heuristics had the highest effectiveness with a score of 0.79. Berry’s heuristics came next with a score of 0.62. Nielsen’s heuristics had the lowest overall effectiveness with a score of 0.31.
5.4.6 Reliability – Differences
Recall that the reliability of each heuristic set is measured in two ways: one relying upon the actual differences among the evaluators, the other upon the average number of agreements among the evaluators.
Here we focus on the former. For example, Berry’s set has eight heuristics, so consider calculating the differences in the ratings for the first heuristic for the first claim in the Plasma Poster. This difference is found by subtracting the ratings of each evaluator from every other evaluator and summing up each of the differences, then dividing by the number of differences (or the average difference). Suppose that an evaluator rated the first heuristic with a 6 (agree) and another rated it as a 4 (neutral) and a third rated it as a 5 (somewhat agree). The difference in this is 1.33.
We then averaged the differences for every heuristic on a given claim to get an overall difference score for that claim, with a lower score indicating higher reliability (zero difference indicates complete reliability). These average differences provide a measure for the reliability of the heuristic set.
Overall Reliability Differences
Considering all 33 claims across the three systems gives an overall indication of the average differences for the heuristic sets. One-way ANOVA suggests significant differences among the three heuristic sets (F(2, 23) = 23.02,MSE = 0.84, p < 0.05).
Pair-wise t-tests show that Somervell’s heuristics had significantly lower average differences than both Berry’s heuristics (df = 14, t = 4.3, p < 0.05) and Nielsen’s heuristics (df = 16, t = 6.8, p < 0.05). No significant differences were found between Berry’s heuristics and Nielsen’s heuristics (df = 16, t = 1.43, p = 0.17), but Berry’s set had a slightly lower average difference (MB = 2.02, SDB = 0.21; MN = 2.14, SDN = 0.13).
5.4.7 Reliability – Agreement
In addition to the average differences, a further measure of reliability was calculated by counting the number of agreements among the evaluators, then dividing by the total number of possible agreements. This calculation provides a measure of the agreement rating for each heuristic.
For example, consider the previous three evaluators and their ratings (6, 5, and 4). The agreement rating in this case would be:
agreement = 0 / 3 = 0
because none of the evaluators agreed on the rating, but there were potentially three agreements (if they had all given the same rating). Averages across all of the claims for a given system were then taken. This provides an assessment of the average agreement for each heuristic as it pertains to a given system.
Overall Agreement
Taking all 33 claims into consideration, one-way ANOVA indicates significant differences among the three heuristic sets for evaluator agreement (F(2, 23) = 6.31,MSE = 0.01, p = 0.01). Pairwise t-tests show that both Somervell’s heuristics and Berry’s heuristics had significantly higher agreement than Nielsen’s set (df = 16, t = 2.99, p = 0.01 and df = 16, t = 3.7, p < 0.05 respectively). No significant differences were found between Somervell’s and Berry’s heuristics (df = 14, t = 0.46, p = 0.65).
5.4.8 Time Spent
Recall that we also asked the evaluators to report the amount of time (in minutes) they spent completing this evaluation. This measure is valuable in assessing the cost of the methods in terms of effort required. It was anticipated that the time required for each method would be similar across the methods.
Averaging reported times across evaluators for each method suggests that Somervell’s set required the least amount of time (M = 103.17, SD = 27.07), but one-way ANOVA reveals no significant differences (F(2, 17) = 0.26, p = 0.77). Berry’s set required the most time (M = 119.14, SD = 60.69) while Nielsen’s set (M = 104.29, SD = 38.56) required slightly more than Somervell’s.
5.5 Discussion
So what does all this statistical analysis mean? What do we know about the three heuristic sets? How have we supported or refuted our hypotheses through this analysis?
5.5.1 Hypotheses Revisited
1. Somervell’s set of heuristics will have a higher validity score for the Notification Collage.
2. More specific heuristics will have higher thoroughness, validity, and reliability measures.
3. Generic methods will require more time for evaluators to complete the study.
Hypothesis 1
For hypothesis one, we discovered that Somervell’s heuristics indeed held the highest validity score for the Notification Collage (see Figure 5.7).
However, this validity score was not 100%, as was expected. What does this mean? It simply illustrates the difference in the evaluators who participated in this study. They did not think that any of the heuristics applied to one of the claims from the Notification Collage. Although, it can be noted that the applicability scores for that particular claims were very close to the cutoff level we chose for agreement (that being 5 or greater on a 7-point scale).
Still, evidence suggests that hypothesis 1 holds.
Hypothesis 2
We find evidence to support this hypothesis based on the scores on each of the three measures: thoroughness, validity, and reliability. In each case, more specific methods had the better ratings over Nielsen’s heuristics for each measure.
Overall one could argue that Somervell’s set of heuristics is most suitable for evaluating large screen information exhibits, but must concede that Berry’s heuristics could also be used with some effectiveness.
Hypothesis 3
We did not find evidence to support this hypothesis. As reported, there were no significant differences in the times required to complete the evaluations for the three methods.
However, Somervell’s and Nielsen’s sets took about 15 fewer minutes, on average, to complete. This does not indicate that the more generic method (Nielsen’s) required more time.
So what would cause the evaluators to take more time with Berry’s method? Initial speculation would suggest that this set uses terminology associated with Notification Systems [62] (see Figure 5.2 for listing of heuristics), including reference to the critical parameters of interruption, reaction, and comprehension, and thus could have increased the interpretation time required to understand each of the heuristics.
5.6 Summary
We have described an experiment to compare three sets of heuristics, representing different levels of generality/specificity, in their ability to evaluate three different LSIE systems. Information on the systems used, test setup, and data collection and analysis has been provided. This test was performed to illustrate the utility that system-class specific methods provide by showing how they are better suited to evaluation of interfaces from that class.
In addition, this work has provided important validation of the creation method used in developing these new heuristics.
We have shown that a system-class specific set of heuristics provides better thoroughness, validity, and reliability than more generic sets (like Nielsen’s). The implication being that without great effort to tailor these generic evaluation tools, they do not provide as effective usability data as a more specific tool.
Source: Somervell, Jacob. Developing Heuristic Evaluation Methods for Large Screen Information Exhibits Based on Critical Parameters. [Dissertation, PhD in Computer Science and Applications] Virginia Polytechnic Institute and State University. June 22, 2004.
Heuristic Comparison Experiment
5.1 Introduction
Now that there is a set of heuristics tailored for the large screen information exhibit system class, a comparison of this set to more established types of heuristics can be done. The purpose of this comparison would be to show the utility of this new heuristic set. This comparison needs to be fair, so that determining the effectiveness of the new method will be accurate.
To assess whether the new set of heuristics provides better usability results over existing alternative sets, we conducted a comparison experiment in which each of three sets of heuristics were used to evaluate three separate large screen information exhibits. We then compared the results of each set through several metrics to determine the better evaluation methods for large screen information exhibits.
5.2 Approach
The following sections provide descriptions of the heuristics used, the comparison method, and the systems used in this experiment.
5.2.1 Heuristic Sets
We used three different sets of usability heuristics, each at a different level of specificity for application to large screen information exhibits, ranging from a set completely designed for this particular system class, to a generic set applicable to a wide range of interactive systems.
Nielsen
The least specific set of heuristics was taken from Nielsen and Mack [70]. This set is intended for use on any interactive system, mostly targeted towards desktop applications. Furthermore, this set has been in use since around 1990. It has been tested and criticized for years, but still remains popular with usability practitioners. Again, this set is not tailored for large screen information exhibits in any way and has no relation to the critical parameters for notification systems.
Visibility of system status
Match between system and real world
User control and freedom
Consistency and standards
Error prevention
Recognition rather than recall
Flexibility and efficiency of use
Aesthetic and minimalist design
Help users recognize, diagnose, and recover from errors
Help and documentation
Figure 5.1: Nielsen’s heuristics. General heuristics that apply to most interfaces. Found in [70].
Berry
The second heuristic set used in this comparison test was created for general notification systems by Berry [9]. This set is based on the critical parameters associated with notification systems [62], but only at cursory levels. This set is more closely tied to large screen information exhibits than Nielsen’s method in that large screen information exhibits are a subset of notification systems, but this set is still generic in nature with regards to the specifics surrounding the LSIE system class.
Notifications should be timely
Notifications should be reliable
Notification displays should be consistent (within priority levels)
Information should be clearly understandable by the user
Allow for shortcuts to more information
Indicate status of notification system
Flexibility and efficiency of use
Provide context of notifications
Allow adjustment of notification parameters to fit user goals
Figure 5.2: Berry’s heuristics. Tailored more towards Notification Systems in general. Found in
[9].
Somervell
The final heuristic set is the one created in this work, as reported in Chapter 4. This set is tailored specifically to large screen information exhibits, and thus would be the most specific method of the three when targeting this type of system. It is based on specific levels of the critical parameters associated with the LSIE system class.
5.2.2 Comparison Technique
To determine which of the three sets is better suited for formative evaluation of large screen information exhibits, we use a current set of comparison metrics that rely upon several measures of a method’s ability to uncover usability problems through an evaluation.
The comparison method we are using typically relies on five separate measures to assess the
utility of a given UEM for one’s particular needs, but we will only use a subset in this particular
comparison study.
Hartson et al. report that thoroughness, validity, effectiveness, reliability, and
downstream utility are appropriate measures for comparing evaluation methods [40].
Specifically, our comparison method capitalizes on thoroughness, validity, effectiveness, and reliability, abandoning the downstream utility measure. This choice is used here because longterm studies are required to illustrate downstream utility.
Thoroughness
This measure gives an indication of a method’s ability to uncover a significant percentage of the problems in a given system. Thoroughness consists of a simple calculation of the number of problems uncovered by a single UEM divided by the total number of problems found by all three methods.
thoroughness = (# of problems found by target UEM) / (# of problems found by all methods)
Validity
Validity refers to the ability of a method to uncover the types of problems that real users would experience in day to day use of the system, as opposed to simple or minor problems. Validity is measured as the number of real problems found divided by the total number of real problems identified in the system.
validity = (# of problems found by target UEM) / (# of problems in the system)
The number of real problems in the system refers to the problem set identified through some standard method that is separate from the method being tested.
Effectiveness
Effectiveness combines the previous two metrics into a single assessment of the method. This measure is calculated by multiplying the thoroughness score by the validity score.
effectiveness = thoroughness X validity
Reliability
Reliability is a measure of the consistency of the results of several evaluators using the method. This is also sometimes referred to as inter-rater reliability. This measure is taken more as agreement between the usability problem sets produced by different people using a given method.
This measure is calculated from the differences among all of the evaluators for a specific system as well as by the total number of agreements among the evaluators, thus two measures are used to provide a more robust measurement of the reliability of the heuristic sets:
reliability-d = difference among evaluators for a specific method
reliability-a = average agreement among evaluators for a specific method
For calculating reliability, Hartson et al. recommend using a method from Sears [81] that depends on the ratio of the standard deviation of the numbers of problems found by the average number found [40]. This measure of reliability is overly complicated for current needs, thus a more traditional measure that relies upon actual rater differences is used instead.
5.2.3 Systems
Three systems were used in the comparison study providing a range of applications for which each heuristic would be used in an analytic evaluation. The intent was to provide enough variability in the test systems to tease out differences in the methods.
1 Source Viewer
2 Plasma Poster
3 Notification Collage
Why Source Viewer?
The Source Viewer was chosen as a target system for this study because we wanted an example of a real system that has been in regular use for an extended period. We immediately thought of command and control situations. Potential candidates included local television stations, local air traffic control towers, electrical power companies, and telephone exchange stations. We finally settled on local television command and control after limited responses from the other candidates.
Why Plasma Poster?
We wanted to include the Plasma Poster because it is one of very few LSIE systems that has seen some success in terms of long term usage and acceptance. It has seen over a year of deployment in a large research laboratory, with reports on usage and user feedback reported in [20].
This lengthy deployment and data collection period provides ample evidence for typical usability problems. We can use the published reports as support for our problem sets. Coupled with developer feedback, we can effectively validate the problem set for this system.
Why Notification Collage?
We chose the Notification Collage as the third system for several reasons.
First, we wanted to increase the validity of any results we find. By using more systems, we get a better picture of the “goodness” of the heuristic sets, especially if we get consistent results across all three systems.
Secondly, we wanted to explicitly show that the heuristic set we created in this work actually uncovered the issues that went into that creation process. In other words, since the Notification Collage was one of the systems that led to this heuristic set, using that set on the Notification Collage should uncover most of the issues with that system.
Finally, we wanted to use the Notification Collage out of the original five because we had the most developer feedback on that system, and like the Plasma Poster, it has seen reasonable deployment and use.
5.2.4 Hypotheses
We have three main hypotheses to test in this experiment:
1. Somervell’s set of heuristics has a higher validity score for the Notification Collage.
We believed this was true because the Notification Collage was used in the creation of Somervell’s heuristics, thus those heuristics should identify most or all of the issues in the Notification Collage.
2. More specific heuristics have higher thoroughness, validity, and reliability measures.
We felt this was true because more specific methods are more closely related to the systems in this study. Indeed, from Chapter 3, we discussed how previous work suggests system-class level heuristics would be best. This experiment illustrates this case for heuristic evaluation of large screen information exhibits.
3. Generic methods require more time for evaluators to complete the study.
This seems logical because a more generic heuristic set would require more interpretation and thought, hence we felt that those evaluators who use Nielsen’s set would take longer to complete the system evaluations, providing further impetus for developing system-class UEMs.
5.2.5 Identifying Problem Sets
One problem identified in other UEM comparison studies involves the calculation of specific metrics that rely upon something referred to as the “real” problem set (see [40]). In most cases, this problem set is the union of the problems found by each of the methods in the comparison study. In other words, each UEM is applied in a standard usability evaluation of a system, and the “real” problem set is simply the union of the problems found by each of the methods.
This comparison study also faced the same challenge. Instead of relying on evaluators to produce sets of problems from each method, then using the union of those problem sets as the “real” problem set, analysis and testing was performed on the target systems beforehand and the problem reports from those efforts were used to come up with a standard set of problems for each system.
Source Viewer Problem Set
To determine the problem sets experienced by the users of this system, a field study was conducted. Two interviews with the users of the large screen system, as well as observation were conducted.
Plasma Poster Problem Set
Analytic evaluation augmented with developer feedback and literature review served as the method for determining the real problem set for the Plasma Poster. We employed the same claims analysis technique that we used in the creation process to identify typical usability tradeoffs for the Plasma Poster. After identifying the usability issues, we asked the developers of the system to verify the tradeoffs.
Notification Collage Problem Set
To validate the problem set for the Notification Collage, we contacted the developers of the system and asked them to check each tradeoff as it pertained to the behavior of real users. The developers were given a list of the tradeoffs found in our claims analysis (from Chapter 4) and asked to verify each tradeoff according to their observations of real user behavior.
5.3 Testing Methodology
This experiment involves a 3x3 mixed factors design. We have three levels of heuristics (Nielsen, Berry, and Somervell) and three systems (Source Viewer, Plasma Poster, and Notification Collage).
The heuristics variable is a between-subjects variable because each evaluator sees only one set of heuristics. The system variable is within-subjects because each participant sees all three systems. For example, evaluator 1 saw only Nielsen’s heuristics, but used those to evaluate all three systems.
We used a balanced Latin Square to ensure learning effects from system presentation order would be minimized. Thus, we needed a minimum of 18 participants (6 per heuristic set) to ensure coverage of the systems in the Latin Square balancing.
5.3.1 Participants
As shown in Table 5.1, we needed a minimum of 18 evaluators for this study.
Twenty-one computer science graduate students who had completed a course on Usability Engineering volunteered for participation as inspectors. Six participants were assigned to each heuristic set, to cover each of the order assignments. Three additional students volunteered and they were randomly assigned a presentation order.
These participants all had knowledge of usability evaluation, as well as analytic and empirical methods. Furthermore, each was familiar with heuristic evaluation. Some of the participants were not familiar with the claim structure used in this study, but they were able to understand the tradeoff concept immediately.
Unfortunately, one of the participants failed to complete the experiment. This individual apparently decided the effort required to complete the test was too much, and thus filled out the questionnaire using a set pattern.
This makes the final number of participants 20, with seven for Nielsen’s heuristics, seven for Berry’s heuristics, and six for Somervell’s heuristics.
5.3.2 Materials
Each target system was described in one to three short scenarios, and screen shots were provided to the evaluators. The goal was to provide the evaluators with a sense of the display and its intended usage.
This material is sufficient for the heuristic inspection technique according to Nielsen and Mack [70]. This setup ensured that each of the heuristic sets would be used with the same material, thereby reducing the number of random variables in the execution of this experiment.
A description of the heuristic set to be used was also provided to the evaluators. This description included a listing of the heuristics and accompanying text clarification. This clarification helps a person understand the intent and meaning of a specific heuristic, hopefully aiding in assessment. These descriptions were taken from [70] and [9] for Nielsen and Berry respectively.
Armed with the materials for the experiment, the evaluator then proceeded to rate each of the heuristics using a 7-point Likert scale, based on whether or not they felt that the heuristic applied to a claim describing a design tradeoff in the interface. Thus they are judging whether or not a specific heuristic applies to the claim, and how much so.
Marks of four or higher indicate agreement that the heuristic applies to the claim, otherwise the evaluator is indicating disagreement that the heuristic applies.
5.3.3 Questionnaire
As mentioned earlier, the evaluators in this experiment provided their feedback through a Likert scale, with agreement ratings for each of the heuristics in the set. In addition to this feedback, each evaluator also rated the claim in terms of how much they felt it actually applied to the interface in question.
By indicating their agreement level with the claim to the interface, we get feedback on whether usability experts actually think the claim is appropriate for the interface in question.
After rating each of the heuristics for the claims, we also asked each evaluator to rank the severity that the claim would hold, if the claim were indeed a usability problem in the interface.
5.3.4 Measurements Recorded
The data collected in this experiment consists of each evaluator’s rating of the claim applicability, each heuristic rating for an individual claim, and the evaluator’s assessment of the severity of the usability problem. This data was collected for each of the thirty claims across the three systems.
In addition to the above measures, we also collected data on the evaluator’s experience with usability evaluation, heuristics, and large screen information exhibits. This evaluator information was collected through survey questions before the evaluation was started.
After the evaluators completed the test, they recorded the amount of time they spent on the task. This was a self reported value as each evaluator worked at his/her own pace and in their own location.
5.4 Results
Twenty-one evaluators provided feedback on 33 different claims across three systems. Each evaluator ended up providing either 10 or 12 question responses per claim, depending on the heuristic set used (Nielsen’s set has 10 in it, whereas the others only have 8). This means we have either 330 or 396 answers to consider, per evaluator.
Fortunately, this data was separable into manageable chunks, dealing with applicability, severity, and heuristic ratings; as well as evaluator experience levels and time to complete for each method.
5.4.1 Participant Experience
As for individual evaluator abilities, the average experience level with usability evaluation, across all three systems, was “amateur”. This means that overall, for each heuristic set, we had comparable experience for the evaluators assigned to that set.
5.4.2 Applicability Scores
To indicate whether or not a heuristic set applied to a given claim (or problem), evaluators marked their agreement with the statement “the heuristic applies to the claim”. This agreement rating indicates that a specific heuristic applied to the claim. Each of the heuristics was marked on a 7-point Likert scale by the evaluators, indicating his/her level of agreement with the statement.
Using this applicability measure, the responses were averaged for a single claim across all of the evaluators. Averaging across evaluators allows assessment of the overall “applicability” of the heuristic to the claim.
This applicability score is used to determine whether any of the heuristics applied to the issue described in the claim. If a heuristic received an “agree” rating, average greater than or equal to five, then that heuristic was thought to have applied to the issue in the claim.
Overall Applicability
Considering all 33 claims together (found in all three systems), one-way analysis of variance (ANOVA) indicates significant differences among the three heuristic sets for applicability (F2,855) = 3.0,MSE = 49.7, p < 0.05). Further pair-wise t-tests reveal that Somervell’s set of heuristics had significantly higher applicability ratings over both Berry’s (df = 526, t = 3.32, p < 0.05) and Nielsen’s sets (df = 592, t = 11.56, p < 0.05). In addition, Berry’s heuristics had significantly higher applicability scores over Nielsen’s set (df = 592, t = 5.94, p < 0.05).
5.4.3 Thoroughness
Recall that thoroughness is measured as the number of problems found by a single method, divided by the number uncovered by all of the methods. This requires a breakdown of the total number of claims into the numbers for each system.
Plasma Poster has 14 claims, Notification Collage has eight claims, and the Source Viewer has 11 claims. We look at thoroughness measures for each system. To calculate the thoroughness measures for the data we have collected, we count the number of claims “covered” by the target heuristic set.
Here we are defining covered to mean that at least one of the heuristics in the set had an average agreement rating of at least five. Why five?
On the Likert scale, five indicates somewhat agree. If we require that the average score across all of the evaluators to be greater than or equal to five for a single heuristic, we are only capturing those heuristics that truly apply to the claim in question.
Overall Thoroughness
Across all three heuristic sets, 28 of 33 claims had applicability scores higher than five. Somervell’s heuristics had the highest thoroughness rating of the three heuristic sets with 96% (27 of 28 claims). Berry’s heuristics came next with a thoroughness score of 86% (24 of 28) and Nielsen’s heuristics had a score of 61 (17 of 28).
5.4.4 Validity
Validity measures the UEM’s ability to uncover real usability problems in a system [40].
Here the full set of problems in the system is used as the real problem set (as discussed in earlier sections).
As with thoroughness, the applicability scores determine the validity each heuristic set held for the three systems. As before, we used the cutoff value of five on the Likert scale to indicate applicability of the heuristic to the claim. An average rating of five or higher indicates that the heuristic applied to the claim in question.
Overall Validity
Similar to thoroughness, validity scores were calculated across all three systems. Out of 33 total claims, only 28 showed applicability scores greater than five across all three heuristic sets. Somervell’s heuristics had the highest validity, with 27 of 33 claims yielding applicability scores greater than five, for a validity score of 82%. Berry’s heuristics had the next highest validity with 24 of 33 claims, for a validity score of 73%. Nielsen’s heuristics had the lowest validity score, with 17 of 33 claims for a score of 52%.
5.4.5 Effectiveness
Effectiveness is calculated by multiplying thoroughness by validity. UEMs that have high thoroughness and high validity will have high effectiveness scores. A low score on either of these measures will reduce the effectiveness score.
Overall Effectiveness
Considering the effectiveness scores across all three systems reveals that Somervell’s heuristics had the highest effectiveness with a score of 0.79. Berry’s heuristics came next with a score of 0.62. Nielsen’s heuristics had the lowest overall effectiveness with a score of 0.31.
5.4.6 Reliability – Differences
Recall that the reliability of each heuristic set is measured in two ways: one relying upon the actual differences among the evaluators, the other upon the average number of agreements among the evaluators.
Here we focus on the former. For example, Berry’s set has eight heuristics, so consider calculating the differences in the ratings for the first heuristic for the first claim in the Plasma Poster. This difference is found by subtracting the ratings of each evaluator from every other evaluator and summing up each of the differences, then dividing by the number of differences (or the average difference). Suppose that an evaluator rated the first heuristic with a 6 (agree) and another rated it as a 4 (neutral) and a third rated it as a 5 (somewhat agree). The difference in this is 1.33.
We then averaged the differences for every heuristic on a given claim to get an overall difference score for that claim, with a lower score indicating higher reliability (zero difference indicates complete reliability). These average differences provide a measure for the reliability of the heuristic set.
Overall Reliability Differences
Considering all 33 claims across the three systems gives an overall indication of the average differences for the heuristic sets. One-way ANOVA suggests significant differences among the three heuristic sets (F(2, 23) = 23.02,MSE = 0.84, p < 0.05).
Pair-wise t-tests show that Somervell’s heuristics had significantly lower average differences than both Berry’s heuristics (df = 14, t = 4.3, p < 0.05) and Nielsen’s heuristics (df = 16, t = 6.8, p < 0.05). No significant differences were found between Berry’s heuristics and Nielsen’s heuristics (df = 16, t = 1.43, p = 0.17), but Berry’s set had a slightly lower average difference (MB = 2.02, SDB = 0.21; MN = 2.14, SDN = 0.13).
5.4.7 Reliability – Agreement
In addition to the average differences, a further measure of reliability was calculated by counting the number of agreements among the evaluators, then dividing by the total number of possible agreements. This calculation provides a measure of the agreement rating for each heuristic.
For example, consider the previous three evaluators and their ratings (6, 5, and 4). The agreement rating in this case would be:
agreement = 0 / 3 = 0
because none of the evaluators agreed on the rating, but there were potentially three agreements (if they had all given the same rating). Averages across all of the claims for a given system were then taken. This provides an assessment of the average agreement for each heuristic as it pertains to a given system.
Overall Agreement
Taking all 33 claims into consideration, one-way ANOVA indicates significant differences among the three heuristic sets for evaluator agreement (F(2, 23) = 6.31,MSE = 0.01, p = 0.01). Pairwise t-tests show that both Somervell’s heuristics and Berry’s heuristics had significantly higher agreement than Nielsen’s set (df = 16, t = 2.99, p = 0.01 and df = 16, t = 3.7, p < 0.05 respectively). No significant differences were found between Somervell’s and Berry’s heuristics (df = 14, t = 0.46, p = 0.65).
5.4.8 Time Spent
Recall that we also asked the evaluators to report the amount of time (in minutes) they spent completing this evaluation. This measure is valuable in assessing the cost of the methods in terms of effort required. It was anticipated that the time required for each method would be similar across the methods.
Averaging reported times across evaluators for each method suggests that Somervell’s set required the least amount of time (M = 103.17, SD = 27.07), but one-way ANOVA reveals no significant differences (F(2, 17) = 0.26, p = 0.77). Berry’s set required the most time (M = 119.14, SD = 60.69) while Nielsen’s set (M = 104.29, SD = 38.56) required slightly more than Somervell’s.
5.5 Discussion
So what does all this statistical analysis mean? What do we know about the three heuristic sets? How have we supported or refuted our hypotheses through this analysis?
5.5.1 Hypotheses Revisited
1. Somervell’s set of heuristics will have a higher validity score for the Notification Collage.
2. More specific heuristics will have higher thoroughness, validity, and reliability measures.
3. Generic methods will require more time for evaluators to complete the study.
Hypothesis 1
For hypothesis one, we discovered that Somervell’s heuristics indeed held the highest validity score for the Notification Collage (see Figure 5.7).
However, this validity score was not 100%, as was expected. What does this mean? It simply illustrates the difference in the evaluators who participated in this study. They did not think that any of the heuristics applied to one of the claims from the Notification Collage. Although, it can be noted that the applicability scores for that particular claims were very close to the cutoff level we chose for agreement (that being 5 or greater on a 7-point scale).
Still, evidence suggests that hypothesis 1 holds.
Hypothesis 2
We find evidence to support this hypothesis based on the scores on each of the three measures: thoroughness, validity, and reliability. In each case, more specific methods had the better ratings over Nielsen’s heuristics for each measure.
Overall one could argue that Somervell’s set of heuristics is most suitable for evaluating large screen information exhibits, but must concede that Berry’s heuristics could also be used with some effectiveness.
Hypothesis 3
We did not find evidence to support this hypothesis. As reported, there were no significant differences in the times required to complete the evaluations for the three methods.
However, Somervell’s and Nielsen’s sets took about 15 fewer minutes, on average, to complete. This does not indicate that the more generic method (Nielsen’s) required more time.
So what would cause the evaluators to take more time with Berry’s method? Initial speculation would suggest that this set uses terminology associated with Notification Systems [62] (see Figure 5.2 for listing of heuristics), including reference to the critical parameters of interruption, reaction, and comprehension, and thus could have increased the interpretation time required to understand each of the heuristics.
5.6 Summary
We have described an experiment to compare three sets of heuristics, representing different levels of generality/specificity, in their ability to evaluate three different LSIE systems. Information on the systems used, test setup, and data collection and analysis has been provided. This test was performed to illustrate the utility that system-class specific methods provide by showing how they are better suited to evaluation of interfaces from that class.
In addition, this work has provided important validation of the creation method used in developing these new heuristics.
We have shown that a system-class specific set of heuristics provides better thoroughness, validity, and reliability than more generic sets (like Nielsen’s). The implication being that without great effort to tailor these generic evaluation tools, they do not provide as effective usability data as a more specific tool.
Source: Somervell, Jacob. Developing Heuristic Evaluation Methods for Large Screen Information Exhibits Based on Critical Parameters. [Dissertation, PhD in Computer Science and Applications] Virginia Polytechnic Institute and State University. June 22, 2004.
Labels:
heuristic evaluation,
Somervell,
usability criteria
Nov 4 - Somervell, Heuristics Creation (PhD dissertation)

Chapter 4
Heuristics Creation
4.1 Introduction
Ensuring usability is an ongoing challenge for software developers. Myriad testing techniques exist, leading to a trade-off between implementation cost and results effectiveness.
Usability testing techniques are broken down into analytical and empirical types.
Analytical methods involve inspection of the system, typically experts in the application field, who identify problems in a walkthrough process.
Empirical methods leverage people who could be real users of the application in controlled tests of specific aspects of the system, often to determine efficiency in performing tasks with the system.
Using either type has advantages and disadvantages, but practitioners typically have limited budgets for usability testing. Thus, they need to use techniques that give useful results while not requiring significant funds. Analytic methods fit this requirement more readily for formative evaluation stages.
With the advent of new technologies and non-traditional interfaces, analytic techniques like heuristics hold the key to early and effective interface evaluation.
There are problems with using analytical methods (like heuristics) that can decrease the validity of results [21]. These problems come from applying a small set of guidelines to a wide range of systems, necessitating interpretation of evaluation results. This illustrates how generic guidelines are not readily applicable to all systems [40], and more specific heuristics are necessary.
Our goal was to create a more specific set, tailored to this system class, yet still have a
set that can be generic enough to apply to all systems in this class.
LSIEs focus on very specific user goals based on the critical parameters of interruption, reaction, and comprehension. Differing levels of each parameter (high, medium, or low) define different system classes [62]. We focus on LSIEs which require medium interruption, low to high
reaction, and high comprehension.
4.2 Motivation
Tremendous effort has been devoted to the study of usability evaluation, specifically in comparing analytic to empirical methods.
Nielsen’s heuristics are probably the most notable set of analytical techniques, developed to facilitate formative usability testing [71, 70]. They have come under fire for their claims that heuristic evaluations are comparable to user testing, yet require fewer test subjects. Comparisons of user testing to heuristic evaluation are numerous [48, 50, 90].
Some have worked to develop targeted heuristics for specific application types.
Baker et al. report on adapting heuristic evaluation to groupware systems [5]. They show that applying heuristic evaluation methods to groupware systems is effective and efficient for formative usability evaluation.
Mankoff et al. compare an adapted set of heuristics to Nielsen’s original set [56]. They studied ambient displays with both sets of heuristics and determined that their adapted set is better suited to ambient displays.
4.3 Processes Involved
How does one create a set of heuristics anyway? We could follow the steps of previous researchers and just use pre-existing heuristics, then reason about the target system class, hopefully coming up with a list of new heuristics that prove useful.
Nielsen and Molich explicitly state that the heuristics come from years of experience and reflection. Not surprising as the heuristics emerged some 30 years after graphical interfaces became mainstream. In the case of Nielsen and Mack, they at least validated their method through using it in the analysis of several systems, after they had created their set.
The two mentioned studies relied upon vague descriptions of theoretical underpinnings [5] or simple tweaking of existing heuristics [56].
Our approach to this lack of structure in creating heuristics is to take a logical look at how one might uncover or discover heuristics for a particular type of system. Basically, to gain insight about a certain type of system, one could analyze several example applications in that system class based on the critical parameters for that system class, and then use the results of that analysis to categorize and group the issues discovered into re-usable design guidelines or heuristics.
These stages involve:
• selection of target systems.
• inspection of these systems. An approach like claims analysis [15] provides necessary
structure to knowledge extraction and provides a consistent representation.
• classifying design implications. Leveraging the underlying critical parameters can help
organize the claims found in terms of impacts to those parameters.
• categorizing design implications. Scenario Based Design [77] provides a mechanism for
categorizing design knowledge into manageable parts.
• extracting high level design guidance. Based on the groupings developed in the previous step, high level design guidelines can be formulated in terms of design issues.
• synthesizing potential heuristics. By matching and relating similar issues, heuristics can
be synthesized.
4.4 Selecting Systems
The first step in the creation process requires careful selection of example systems to inspect and analyze for uncovering existing problems in the systems. The idea is to uncover typical issues inherent in that specific type of system.
Our goal was to use a representative set of systems from the LSIE class. We wanted systems that had been in use for a while, with reports on usage or studies on usability to help validate the analysis we would perform on the systems. //we chose the following five LSIE systems, including some from our own work and some from other well-documented design efforts, to further investigate in the creation process.
• GAWK [31] This system provides teachers and students an overview and history of current project work by group and time, on a public display in the classroom.
• Photo News Board [85] This system provides photos of news stories in four categories, shown on a large display in a break room or lab.
• Notification Collage [36] This system provides users with communication information and various data from others in the shared space on a large screen.
• What’s Happening? [94, 95] This system shows relevant information (news, traffic, weather) to members of a local group on a large, wall display.
• BlueBoard [78] This system allows members in a local setting to view information pages about what is occurring in their location (research projects, meetings, events).
These five systems were chosen as a representative set of large screen information exhibits. The GAWK and Photo News Board were created in local labs and thus we have access to the developers and potential user classes. The other three are some of the more famous and familiar ones found in recent literature.
4.5 Analyzing Systems
Now that we have selected our target systems, we must now determine the typical usability issues and problems inherent in these systems. Performing usability analysis or testing of these systems finds the issues and problems each system holds. To find usability problems we can do analytic or empirical investigations, recording the issues we find.
We chose to use an analytic evaluation approach to the five aforementioned LSIEs, based on
arguments from Section 3.3. We wanted to uncover as many usability concerns as possible, so
we chose claims analysis [15, 77] as the analytic vehicle with which we investigated our systems.
4.5.1 Claims Analysis
Claims analysis is a method for determining the impacts design decisions have on user goals for
a specific piece of software [15, 77]. Claims are statements about a design element reflecting a
positive or negative effect resulting from using the design element in a system [15].
Claims analysis involves inspection and reflection on the wordings of specific claims to determine the psychological impacts a design artifact may have on a user [15]. The wordings are the actual words used to describe positive and negative effects of the claims. The impacts are the overall psychological effect on the user.
4.5.2 System Claims
Claims were made for each of the five systems that were inspected. These claims focused on design artifacts and overall goals of the systems. These claims are based on typical usage, as exemplified by the scenarios shown for each system. On average, there were over 50 claims made per system.
Table 4.2 shows the breakdown of the numbers of claims found for each system.
Each claim dealt with some design element in the interface, showing upsides or downsides resulting from a particular design choice.
These claims can be thought of as problem indicators, unveiling potential problems with the system being able to support the user goals. These problem indicators include positive aspects of design choices as well. By including the good with the bad, we gain fuller understanding of the underlying design issues.
4.5.3 Validating Claims
How do we know that the claims we found through our analysis represent the “real” design challenges in the systems? This is a fair question and one that must be addressed. We need to verify that the claims we are using to extract design guidance for LSIE systems are actually representative of real user problems encountered during use of those systems. We tackled this problem through several different techniques.
For the GAWK and Photo News Board, we relied upon existing empirical studies [85] to validate the claims we found for those systems.
For the Notification Collage we relied upon discussion and feedback from the system developers. We sent the list of claims and scenarios to Saul Greenberg and Michael Rounding and asked them to verify that the claims we made for the Notification Collage were typical of what they observed users actually doing with the system. Michael Rounding provided a thorough response that indicated most of the claims were indeed correct and experienced by real users of the system.
A similar effort was attempted with both the What’s Happening? and Blue Board systems. The developers of these systems were contacted but no specific feedback was provided on our claims. However, John Stasko, co-developer of the What’s Happening? system, provided interview feedback on the system and provided a nice publication [95] that served as validation material for the claims. This report provides details on user experiments done with the What’s Happening? system. Using this report, we were able to verify that most of the claims we made for the system were experienced in those experiments.
Unfortunately, none of the developers of the Blue Board system responded to our request. We were able to use existing literature on the system to verify some of the claims but the reports on user behavior in [78] did not provide enough material to validate all of the claims we found for that system.
4.6 Categorizing Claims
Now that we have analyzed several systems in the LSIE class, and we have over 250 claims about design decisions for those systems, how do we make sense of it all and glean reusable design guidance in the form of heuristics? To make sense of the claims we have, we need to group and categorize similar claims.
This requires a framework to ensure consistent classification and facilitate final heuristic synthesis from the classification. This is where the idea of critical parameters plays an important role.
4.6.1 Classifying Claims Using the IRC Framework
Recall that notification systems can be classified by their level of impact on interruption, reaction, and comprehension [62]. This classification scheme can be simplified to reflect a high, medium, or low impact to each of interruption, reaction, and comprehension.
In other words, we can take a single claim and classify it according to the impact it would have on the user goals associated with the system.
For example, we have a claim about the collage metaphor from the Notification Collage system that suggests that the lack of organization can hinder efforts to find information. This claim would be classified as “high” interruption because it increases the time required to find a piece of information. It could also be classified as “low” comprehension because it reduces a person’s ability to understand the information quickly and accurately. It is perfectly acceptable to have the claim fit into both classifications.
4.6.2 Assessing Goal Impact
Determining the impact a claim has on the user goals was done through inspection and reflection techniques. Each claim was read and approached from the scenarios for the system, trying to identify if the claim had an impact on the user goals. A claim impacted a user goal if it was determined through the wording of the claim that one of interruption, reaction, or comprehension was modified by the design element.
To assign user goal impacts to the claims, a team of experts should assess each claim.
These experts should have extensive knowledge of the system class, and the critical parameters that define that class. Knowledge of claims analysis techniques and/or usability evaluation are highly recommended.
We used a two- person team of experts.
Differences occurred when these classifications were not compatible.
Agreement was measured as the number of claims with the same classification divided by the total number of claims. We found that initial agreement on the claims was near 94% and after discussion was 100% for all claims.
This calculation comes from the fact that out of 253 individual claims, 237 were classified by the inspectors as impacting the user goals in the same way, i.e. all of the experts agreed on the same classification.
4.6.3 Categorization Through Scenario Based Design
Heuristics Creation
4.1 Introduction
Ensuring usability is an ongoing challenge for software developers. Myriad testing techniques exist, leading to a trade-off between implementation cost and results effectiveness.
Usability testing techniques are broken down into analytical and empirical types.
Analytical methods involve inspection of the system, typically experts in the application field, who identify problems in a walkthrough process.
Empirical methods leverage people who could be real users of the application in controlled tests of specific aspects of the system, often to determine efficiency in performing tasks with the system.
Using either type has advantages and disadvantages, but practitioners typically have limited budgets for usability testing. Thus, they need to use techniques that give useful results while not requiring significant funds. Analytic methods fit this requirement more readily for formative evaluation stages.
With the advent of new technologies and non-traditional interfaces, analytic techniques like heuristics hold the key to early and effective interface evaluation.
There are problems with using analytical methods (like heuristics) that can decrease the validity of results [21]. These problems come from applying a small set of guidelines to a wide range of systems, necessitating interpretation of evaluation results. This illustrates how generic guidelines are not readily applicable to all systems [40], and more specific heuristics are necessary.
Our goal was to create a more specific set, tailored to this system class, yet still have a
set that can be generic enough to apply to all systems in this class.
LSIEs focus on very specific user goals based on the critical parameters of interruption, reaction, and comprehension. Differing levels of each parameter (high, medium, or low) define different system classes [62]. We focus on LSIEs which require medium interruption, low to high
reaction, and high comprehension.
4.2 Motivation
Tremendous effort has been devoted to the study of usability evaluation, specifically in comparing analytic to empirical methods.
Nielsen’s heuristics are probably the most notable set of analytical techniques, developed to facilitate formative usability testing [71, 70]. They have come under fire for their claims that heuristic evaluations are comparable to user testing, yet require fewer test subjects. Comparisons of user testing to heuristic evaluation are numerous [48, 50, 90].
Some have worked to develop targeted heuristics for specific application types.
Baker et al. report on adapting heuristic evaluation to groupware systems [5]. They show that applying heuristic evaluation methods to groupware systems is effective and efficient for formative usability evaluation.
Mankoff et al. compare an adapted set of heuristics to Nielsen’s original set [56]. They studied ambient displays with both sets of heuristics and determined that their adapted set is better suited to ambient displays.
4.3 Processes Involved
How does one create a set of heuristics anyway? We could follow the steps of previous researchers and just use pre-existing heuristics, then reason about the target system class, hopefully coming up with a list of new heuristics that prove useful.
Nielsen and Molich explicitly state that the heuristics come from years of experience and reflection. Not surprising as the heuristics emerged some 30 years after graphical interfaces became mainstream. In the case of Nielsen and Mack, they at least validated their method through using it in the analysis of several systems, after they had created their set.
The two mentioned studies relied upon vague descriptions of theoretical underpinnings [5] or simple tweaking of existing heuristics [56].
Our approach to this lack of structure in creating heuristics is to take a logical look at how one might uncover or discover heuristics for a particular type of system. Basically, to gain insight about a certain type of system, one could analyze several example applications in that system class based on the critical parameters for that system class, and then use the results of that analysis to categorize and group the issues discovered into re-usable design guidelines or heuristics.
These stages involve:
• selection of target systems.
• inspection of these systems. An approach like claims analysis [15] provides necessary
structure to knowledge extraction and provides a consistent representation.
• classifying design implications. Leveraging the underlying critical parameters can help
organize the claims found in terms of impacts to those parameters.
• categorizing design implications. Scenario Based Design [77] provides a mechanism for
categorizing design knowledge into manageable parts.
• extracting high level design guidance. Based on the groupings developed in the previous step, high level design guidelines can be formulated in terms of design issues.
• synthesizing potential heuristics. By matching and relating similar issues, heuristics can
be synthesized.
4.4 Selecting Systems
The first step in the creation process requires careful selection of example systems to inspect and analyze for uncovering existing problems in the systems. The idea is to uncover typical issues inherent in that specific type of system.
Our goal was to use a representative set of systems from the LSIE class. We wanted systems that had been in use for a while, with reports on usage or studies on usability to help validate the analysis we would perform on the systems. //we chose the following five LSIE systems, including some from our own work and some from other well-documented design efforts, to further investigate in the creation process.
• GAWK [31] This system provides teachers and students an overview and history of current project work by group and time, on a public display in the classroom.
• Photo News Board [85] This system provides photos of news stories in four categories, shown on a large display in a break room or lab.
• Notification Collage [36] This system provides users with communication information and various data from others in the shared space on a large screen.
• What’s Happening? [94, 95] This system shows relevant information (news, traffic, weather) to members of a local group on a large, wall display.
• BlueBoard [78] This system allows members in a local setting to view information pages about what is occurring in their location (research projects, meetings, events).
These five systems were chosen as a representative set of large screen information exhibits. The GAWK and Photo News Board were created in local labs and thus we have access to the developers and potential user classes. The other three are some of the more famous and familiar ones found in recent literature.
4.5 Analyzing Systems
Now that we have selected our target systems, we must now determine the typical usability issues and problems inherent in these systems. Performing usability analysis or testing of these systems finds the issues and problems each system holds. To find usability problems we can do analytic or empirical investigations, recording the issues we find.
We chose to use an analytic evaluation approach to the five aforementioned LSIEs, based on
arguments from Section 3.3. We wanted to uncover as many usability concerns as possible, so
we chose claims analysis [15, 77] as the analytic vehicle with which we investigated our systems.
4.5.1 Claims Analysis
Claims analysis is a method for determining the impacts design decisions have on user goals for
a specific piece of software [15, 77]. Claims are statements about a design element reflecting a
positive or negative effect resulting from using the design element in a system [15].
Claims analysis involves inspection and reflection on the wordings of specific claims to determine the psychological impacts a design artifact may have on a user [15]. The wordings are the actual words used to describe positive and negative effects of the claims. The impacts are the overall psychological effect on the user.
4.5.2 System Claims
Claims were made for each of the five systems that were inspected. These claims focused on design artifacts and overall goals of the systems. These claims are based on typical usage, as exemplified by the scenarios shown for each system. On average, there were over 50 claims made per system.
Table 4.2 shows the breakdown of the numbers of claims found for each system.
Each claim dealt with some design element in the interface, showing upsides or downsides resulting from a particular design choice.
These claims can be thought of as problem indicators, unveiling potential problems with the system being able to support the user goals. These problem indicators include positive aspects of design choices as well. By including the good with the bad, we gain fuller understanding of the underlying design issues.
4.5.3 Validating Claims
How do we know that the claims we found through our analysis represent the “real” design challenges in the systems? This is a fair question and one that must be addressed. We need to verify that the claims we are using to extract design guidance for LSIE systems are actually representative of real user problems encountered during use of those systems. We tackled this problem through several different techniques.
For the GAWK and Photo News Board, we relied upon existing empirical studies [85] to validate the claims we found for those systems.
For the Notification Collage we relied upon discussion and feedback from the system developers. We sent the list of claims and scenarios to Saul Greenberg and Michael Rounding and asked them to verify that the claims we made for the Notification Collage were typical of what they observed users actually doing with the system. Michael Rounding provided a thorough response that indicated most of the claims were indeed correct and experienced by real users of the system.
A similar effort was attempted with both the What’s Happening? and Blue Board systems. The developers of these systems were contacted but no specific feedback was provided on our claims. However, John Stasko, co-developer of the What’s Happening? system, provided interview feedback on the system and provided a nice publication [95] that served as validation material for the claims. This report provides details on user experiments done with the What’s Happening? system. Using this report, we were able to verify that most of the claims we made for the system were experienced in those experiments.
Unfortunately, none of the developers of the Blue Board system responded to our request. We were able to use existing literature on the system to verify some of the claims but the reports on user behavior in [78] did not provide enough material to validate all of the claims we found for that system.
4.6 Categorizing Claims
Now that we have analyzed several systems in the LSIE class, and we have over 250 claims about design decisions for those systems, how do we make sense of it all and glean reusable design guidance in the form of heuristics? To make sense of the claims we have, we need to group and categorize similar claims.
This requires a framework to ensure consistent classification and facilitate final heuristic synthesis from the classification. This is where the idea of critical parameters plays an important role.
4.6.1 Classifying Claims Using the IRC Framework
Recall that notification systems can be classified by their level of impact on interruption, reaction, and comprehension [62]. This classification scheme can be simplified to reflect a high, medium, or low impact to each of interruption, reaction, and comprehension.
In other words, we can take a single claim and classify it according to the impact it would have on the user goals associated with the system.
For example, we have a claim about the collage metaphor from the Notification Collage system that suggests that the lack of organization can hinder efforts to find information. This claim would be classified as “high” interruption because it increases the time required to find a piece of information. It could also be classified as “low” comprehension because it reduces a person’s ability to understand the information quickly and accurately. It is perfectly acceptable to have the claim fit into both classifications.
4.6.2 Assessing Goal Impact
Determining the impact a claim has on the user goals was done through inspection and reflection techniques. Each claim was read and approached from the scenarios for the system, trying to identify if the claim had an impact on the user goals. A claim impacted a user goal if it was determined through the wording of the claim that one of interruption, reaction, or comprehension was modified by the design element.
To assign user goal impacts to the claims, a team of experts should assess each claim.
These experts should have extensive knowledge of the system class, and the critical parameters that define that class. Knowledge of claims analysis techniques and/or usability evaluation are highly recommended.
We used a two- person team of experts.
Differences occurred when these classifications were not compatible.
Agreement was measured as the number of claims with the same classification divided by the total number of claims. We found that initial agreement on the claims was near 94% and after discussion was 100% for all claims.
This calculation comes from the fact that out of 253 individual claims, 237 were classified by the inspectors as impacting the user goals in the same way, i.e. all of the experts agreed on the same classification.
4.6.3 Categorization Through Scenario Based Design
Categorization is needed to separate the claims into manageable groups. By focusing on related claims, similar design tradeoffs can be considered together. An interface design methodology is useful because these approaches often provide a built-in structure that facilitates claims categorization.
Possible design methodologies include Scenario Based Design [77], User Centered Design [73], and Norman’s Stages of Action [72].
Scenario based design (SBD)[77] is an interface design methodology that relies on scenarios
about typical usage of a target system.
Activity Design
Activity design involves what users can and cannot accomplish with the system, at a high level [77]. These are the tasks that the interface supports, ones that the users would otherwise not be able to accomplish.
Activity design encompasses metaphors and supported/unsupported activities [77].
Information Design
Information design deals with how information is shown and how the interface looks [77]. Design decisions for information presentation directly impact comprehension, as well as interruption. Identifying the impacts of information design decisions on user goals can lead to effective design guidelines.
We chose to use the following sub-categories for refining the information design category: use of screen space, foreground and background colors, use of fonts, use of audio, use of animation, and layout. These sub-categories were chosen because they cover almost all of the design issues relevant to information design [77].
Interaction Design
Interaction design focuses on how a user would interact with a system (clicking, typing, etc) [77]. This includes recognizing affordances, understanding the behavior of interface controls, knowing the expected transitions of states in the interface, support for error recovery and undo operations, feedback about task goals, and configurability options for different user classes [77].
Categorization
Armed with the above categories, we are now able to group individual claims into an organized structure, thereby facilitating further analysis and reuse. So how do we know in which area a particular claim should go? This again is done through group analysis and discussion regarding the wording of the claim. The claim wordings typically indicate which category of SBD applies, and any disagreements can be handled through discussion and mitigation.
Similar to the classification effort, this categorization process relied upon the claim wordings for correct placement within the SBD categories. The sub- categories for each of activity, information, and interaction provide 14 areas in which claims may be placed.
Unclassified Claims
Some of the claims were deemed to be unclassified, since the claim did not impact interruption, reaction, or comprehension. While it is possible to situate these claims within the SBD categories, if the claim does not impact one of the three user goals, it was said to be unclassified.
4.7 Synthesis Into Heuristics
After classifying the problems within the framework, we then needed to extract usable design recommendations from those problems.
This required re-inspection of the claim groupings to determine the underlying causes to these issues.
Since the problems come from different systems, we get a broad look at potential design flaws. Identifying and recognizing these flaws in these representative systems can help other designers avoid making those same mistakes in their work.
4.7.1 Visualizing the Problem Tree
To better understand how claims impacted the user goals of each of the systems, a problem tree was created to aid in the visualization of the dispersion of the claims within different areas of the SBD categories.
A problem tree is a collection of claims for a system class, organized by categories, sub-categories, and critical parameter. It serves as a representation of the design knowledge that is collected from the claims analysis process.
A node in the problem tree refers to a collection of claims that fits within a single category (from SBD) with a single classification (from the critical parameters). A leaf in the tree refers to a single claim, and is attached to some node in the tree.
4.7.2 Identifying Issues
To glean reusable design guidance from the individual claims, team discussion was used. A team of experts who are familiar with the claims analysis process and the problem tree considers each node in the tree with the aim of identifying one or more issues that capture the claims within said node.
Issues are design statements, more general than individual claims.
This effort produced 22 issues that covered the 333 claims.
4.7.3 Issues to Heuristics
Armed with the 22 high level issues, we now needed to extract a subset of high level design heuristics from these issues. Twenty-two is unmanageable for formative heuristic evaluation [66] and in many cases the issues were similar or related, suggesting opportunities for concatenation and grouping. This similarity allowed us to create higher level, more generic heuristics to capture the issues.
We created eight final heuristics, capturing the 22 issues discovered in the earlier process.
Table 4.7 provides an example of how we moved from the issues to the heuristics. In most instances, two or three issues could be combined into a single heuristic. However some of the issues were already at a high level and were taken directly into the heuristic list.
4.7.4 Heuristics
Here is the list of heuristics that can be used to guide evaluation of large screen information exhibits.
Explanatory text follows each heuristic, to clarify and illustrate how the heuristics could impact evaluation. Each is general enough to be applied to many systems in this application class, yet they all address the unique user goals of large screen information exhibits.
• Appropriate color schemes should be used for supporting information understanding.
Try using cool colors such as blue or green for background or borders. Use warm colors like red and yellow for highlighting or emphasis.
Try using cool colors such as blue or green for background or borders. Use warm colors like red and yellow for highlighting or emphasis.
• Layout should reflect the information according to its intended use.
Time based information should use a sequential layout; topical information should use categorical, hierarchical, or grid layouts. Screen space should be delegated according to information importance.
• Judicious use of animation is necessary for effective design.
Multiple, separate animations should be avoided. Indicate current and target locations if items are to be automatically moved around the display. Introduce new items with slower, smooth transitions. Highlighting related information is an effective technique for showing relationships among data.
• Use text banners only when necessary.
Reading text on a large screen takes time and effort. Try to keep it at the top or bottom of the screen if necessary. Use sans serif fonts to facilitate reading, and make sure the font sizes are big enough.
• Show the presence of information, but not the details.
Use icons to represent larger information structures, or to provide an overview of the information space, but not the detailed information; viewing information details is better suited to desktop interfaces. The magnitude or density of the information dictates representation mechanism (text vs icons for example).
• Using cyclic displays can be useful, but care must be taken in implementation.
Indicate “where” the display is in the cycle (i.e. 1 of 5 items, or progress bar). Timings (both for
single item presence and total cycle time) on cycles should be appropriate and allow users to
understand content without being distracted.
single item presence and total cycle time) on cycles should be appropriate and allow users to
understand content without being distracted.
• Avoid the use of audio.
Audio is distracting, and on a large public display, could be detrimental to others in the setting. Furthermore, lack of audio can reinforce the idea of relying on the visual system for information exchange.
• Eliminate or hide configurability controls.
Large public displays should be configured one time by an administrator. Allowing multiple users to change settings can increase confusion and distraction caused by the display. Changing the interface too often prevents users from learning the interface.
Source: Somervell, Jacob. Developing Heuristic Evaluation Methods for Large Screen Information Exhibits Based on Critical Parameters. [Dissertation, PhD in Computer Science and Applications] Virginia Polytechnic Institute and State University. June 22, 2004.
Labels:
heuristic evaluation,
Somervell,
usability criteria
Saturday, October 31, 2009
Nov 1 – Somervell’s dissertation
Chapter 2
Literature Review
developing new heuristics for the LSIE system class, based on critical parameters.
Critical Parameters
William Newman put forth the idea of critical parameters for guiding design and strengthening evaluation in [68] as a solution to the growing disparity between interactive system design and separate evaluation.
For example, consider airport terminals, where the critical parameter would be flight capacity per hour per day [68]. All airport terminals can be assessed in terms of this capacity, and improving that capacity would invariably mean we have a better airport. Newman argues that by establishing parameters for application classes, researchers can begin establishing evaluation criteria, thereby providing continuity in evaluation that allows us “to tell whether progress is being made” [68].
In addition, Newman argues that critical parameters can actually provide support for developing design methodologies, based on the most important aspects of a design space. This ability separates critical parameters from traditional usability metrics. Most usability metrics, like “learnability” or “ease of use” only probe the interaction of the user with some interface, focusing not on the intended purpose of the system but on what the user can do with the system.
Critical parameters focus on supporting the underlying system functions that allow one to determine whether the system performs its intended tasks.
Indeed, the connection between critical parameters and traditional usability metrics can be described as input and output of a “usability” function. Critical parameters are used to derive the appropriate usability metrics for a given system, and these metrics are related to the underlying system goals through the critical parameters.
Critical Paramters for Notification Systems
In [62], we embraced Newman’s view of critical parameters and established three parameters that define the notification systems design space.
Interruption, reaction, and comprehension are three attributes of all notification systems that allow one to assess whether the system serves its intended use.
Furthermore, these parameters allow us to assess the user models and system designs associated with notification systems in terms of how well a system supports these three parameters.
High and low values of each parameter capture the intent of the system, and allow one to measure whether the system supports these intents.
2.2.1 Analytical Methods
Analytical methods show great promise for ensuring formative evaluation is completed, and not just acknowledged in the software life cycle. These methods provide efficient and effective usability results [70].
The alternative usually involves costly user studies, which are difficult to perform, and increase the design phases for most interface development projects. It is for these reasons that we focus on analytical methods, specifically heuristics.
Heuristic methods are chosen in this research for two reasons.
One, these methods are considered “discount” methods because they require minimal resources for the usability problems they uncover [70].
Two, these methods only require system mock-ups or screen shots for evaluation, which makes them desirable for formative evaluation. These are strong arguments for developing this method for application in multiple areas.
2.2.2 Heuristic Evaluation
A popular evaluation method, both in academia and industry is heuristic evaluation.
Heuristics are simple, fast approaches to assessing usability [70]. Expert evaluators visually inspect an interface to determine problems related to a set of guidelines (heuristics). These experts identify problems based on whether or not the interface fails to adhere to a given heuristic. When there is a failure, there is typically a usability problem. Studies of heuristics have shown them to be effective (in terms of numbers of problems found) and efficient (in terms of cost to perform) [48, 50].
Some researchers have illustrated difficulties with heuristic evaluation. Cockton & Woolrych suggest that heuristics should be used less in evaluation, in favor of empirical evaluations involving users [21]. Their arguments revolve around discrepancies among different evaluators and the low number of major problems that are found through the technique. Gray & Salzman also point out this weakness in [32].
Despite these objections, heuristic evaluation methods, particularly Nielsen’s, are still popular for their “discount” [70] approach to usability evaluation. Several recent works deal with adapting heuristic approaches to specified areas.
Baker et al. report on adapting heuristic evaluation to groupware systems [5]. They show that applying heuristic evaluation methods to groupware systems is effective and efficient for formative usability evaluation.
Mankoff et al. actually compare an adapted set of heuristics to Nielsen’s original set [56]. They studied ambient displays (which are similar to the systems that would be classified as ambient displays in the IRC framework) with both sets of heuristics and determined that their adapted set is better suited to ambient displays.
The heuristic usability evaluation method will be investigated in this research, but with different forms of heuristics, some adapted specifically to large screen information exhibits, others geared towards more general interface types (like generic notification systems or simply interfaces).
The focus of our work is to create a new set of heuristics by reliance on critical parameters. One that is tailored to the LSIE system class.
2.2.3 Comparing UEMs
Recent examples of work that strives to compare heuristic approaches to other UEMs (like lab-based user testing) include work shown at the 46th Annual Meeting of the Human Factors and Ergonomics Society.
Chattratichart and Brodie report on a comparison study of heuristic methods [16]. They extended heuristic evaluation (based on Nielsen’s) with a small set of content areas. These content areas served to focus the evaluation, thus producing more reliable results. It should also be noted that subjective opinions about the new method favored the original approach over the new approach. The added complexity of grouping problems into the content areas is the speculated cause of this finding [16].
Tan and Bishu compared heuristic evaluation to user testing [90]. They focused their work on web page evaluation and found that heuristic evaluation found more problems, but that the two techniques found different classes of problems. This means that these two methods are difficult to compare since the resulting problem lists are so different (like comparing apples to oranges). This difficulty in comparing analytical to empirical methods has been debated (see Human Computer Interaction 13(4) for a great summary of this debate) before and this particular work brings it to light in a more current example.
There has been some work on the best ways to compare UEMs. These studies are often limited to a specific area within HCI.
For example, Lavery et al. compared heuristics and task analysis in the domain of software visualization [52]. Their work resulted in development of problem reports that facilitate comparison of problems found with different methods. Their comparisons relied on effectiveness, efficiency, and validity measures for each method.
Others have also pointed out that effectiveness, efficiency, and validity are desirable measures for comparing UEMs (beyond simple numbers of usability problems obtained through the method) [40, 21].
Hartson et al. further put forth thoroughness, validity, reliability, and downstream utility as measures for comparing usability evaluation methods [40].
Chapter 3
Background and Motivation
However, there are many different types of usability evaluation methods one could employ to test design, and it is unclear which ones would serve as the best for this system class (large screen information exhibits).
One important variation in methods is whether to use an interface-specific tool or a generic tool that applies to a broad class of systems.
This preliminary study investigates tradeoffs of these two approaches (generic or specific) for evaluating LSIEs, by applying two types of evaluation to example LSIE systems.
This work provides the motivation and direction for the creation, testing, and use of a new set of heuristics tailored to the LSIE system class.
3.2 Assessing Evaluation Methods
Specific evaluation tools are developed for a single application, and apply solely to the system being tested (we refer to this as a per-study basis).
Many researchers use this approach, creating evaluation metrics, heuristics, or questionnaires tailored to the system in question (for example see [5, 56]). These tools seem advantageous because they provide fine grained insight into the target system, yielding detailed redesign solutions. However, filling immediate needs is costly—for each system to be tested a new evaluation method needs to be designed (by designers or evaluators), implemented, and used in the evaluation phase of software development.
In contrast, system-class evaluation tools are not tailored to a specific system and tend to focus on higher level, critical problem areas that might occur in systems within a common class.
These methods are created once (by usability experts) and used many times in separate evaluations. They are desirable for allowing ready application, promoting comparison between different systems, benchmarking system performance measures, and recognizing long-term, multi-project development progress.
However, using a system-class tool often means evaluators sacrifice focus on important interface details, since not all of the system aspects may be addressed by a generic tool. The appeal of system- class methods is apparent over a long-term period, namely through low cost and high benefit.
We conducted an experiment to determine the benefits of each approach in supporting a claims analysis, a key process within the scenario-based design approach [15, 77]. In a claims analysis, an evaluator makes claims about how important interface features will impact users.
Claims can be expressed as tradeoffs, conveying upsides or downsides of interface aspects like supported or unsupported activities, use of metaphors, information design choices (use of color, audio, icons, etc.), or interaction design techniques (affordances, feedback, configuration options, etc.). These claims capture the psychological impacts specific design decisions may have on users.
3.3 Motivation from Prior Work
UEM research efforts have developed high level, generic evaluation procedures, a notable example being Nielsen’s heuristics [70].
Heuristic evaluation has been embraced by practitioners because of its discount approach to assessing usability. With this approach (which involves identification of usability problems that fall into nine general and “most common problem areas”), 3-5 expert evaluators can uncover 70% of an interface’s usability problems.
However, the drawbacks to this approach (and most generic approaches) are evident in the need to develop more specific versions of heuristics for particular classes of systems.
For example, Mankoff et al. created a modified set of heuristics for ambient displays [56]. These displays differ from regular interfaces in that they often reside off the desktop, incorporating parts of the physical space in their design, hence necessitating a more specific approach to evaluation. They came up with the new set of heuristics by eliminating some from Nielsen’s original set, modifying the remaining heuristics to reflect ambient wording, and then added five new heuristics [56]. However, they do not report the criteria used in eliminating the original heuristics, the reasons for using the new wordings, or how they came up with the five new heuristics. They proceeded to compare this new set of heuristics to Nielsen’s original set and found the more specific heuristics provided better usability results.
Similar UEM work dealt with creating modified heuristics for groupware systems [5]. In this work, Baker et al. modified Nielsen’s original set to more closely match the user goals and needs associated with groupware systems. They based their modification on prior groupware system models to provide guidance in modifying Nielsen’s heuristics. The Locales Framework [35] and the mechanics of collaboration [38] helped Baker et al. in formulating their new heuristics. However, they do not describe how these models helped them in their creation, nor how they were used. From the comparison, they found the more application class-specific set of heuristics produced better results compared to the general set (Nielsen’s).
Both of these studies suggest that system-class specific heuristics are more desirable for formative evaluation. However, the creation processes used in both are not adequately described. It seems that to obtain the new set of heuristics, all the researchers did was modify Nielsen’s heuristics.
Unfortunately, it is not clear how this modification occurred. Did the researchers base the changes on important user goals for the system, as determined through critical parameters for the system class? Or was the modification based on guesswork or simple “this seems important for this type of system” style logic?
Source: Somervell, Jacob. Developing Heuristic Evaluation Methods for Large Screen Information Exhibits Based on Critical Parameters. [Dissertation, PhD in Computer Science and Applications] Virginia Polytechnic Institute and State University. June 22, 2004.
Literature Review
developing new heuristics for the LSIE system class, based on critical parameters.
Critical Parameters
William Newman put forth the idea of critical parameters for guiding design and strengthening evaluation in [68] as a solution to the growing disparity between interactive system design and separate evaluation.
For example, consider airport terminals, where the critical parameter would be flight capacity per hour per day [68]. All airport terminals can be assessed in terms of this capacity, and improving that capacity would invariably mean we have a better airport. Newman argues that by establishing parameters for application classes, researchers can begin establishing evaluation criteria, thereby providing continuity in evaluation that allows us “to tell whether progress is being made” [68].
In addition, Newman argues that critical parameters can actually provide support for developing design methodologies, based on the most important aspects of a design space. This ability separates critical parameters from traditional usability metrics. Most usability metrics, like “learnability” or “ease of use” only probe the interaction of the user with some interface, focusing not on the intended purpose of the system but on what the user can do with the system.
Critical parameters focus on supporting the underlying system functions that allow one to determine whether the system performs its intended tasks.
Indeed, the connection between critical parameters and traditional usability metrics can be described as input and output of a “usability” function. Critical parameters are used to derive the appropriate usability metrics for a given system, and these metrics are related to the underlying system goals through the critical parameters.
Critical Paramters for Notification Systems
In [62], we embraced Newman’s view of critical parameters and established three parameters that define the notification systems design space.
Interruption, reaction, and comprehension are three attributes of all notification systems that allow one to assess whether the system serves its intended use.
Furthermore, these parameters allow us to assess the user models and system designs associated with notification systems in terms of how well a system supports these three parameters.
High and low values of each parameter capture the intent of the system, and allow one to measure whether the system supports these intents.
2.2.1 Analytical Methods
Analytical methods show great promise for ensuring formative evaluation is completed, and not just acknowledged in the software life cycle. These methods provide efficient and effective usability results [70].
The alternative usually involves costly user studies, which are difficult to perform, and increase the design phases for most interface development projects. It is for these reasons that we focus on analytical methods, specifically heuristics.
Heuristic methods are chosen in this research for two reasons.
One, these methods are considered “discount” methods because they require minimal resources for the usability problems they uncover [70].
Two, these methods only require system mock-ups or screen shots for evaluation, which makes them desirable for formative evaluation. These are strong arguments for developing this method for application in multiple areas.
2.2.2 Heuristic Evaluation
A popular evaluation method, both in academia and industry is heuristic evaluation.
Heuristics are simple, fast approaches to assessing usability [70]. Expert evaluators visually inspect an interface to determine problems related to a set of guidelines (heuristics). These experts identify problems based on whether or not the interface fails to adhere to a given heuristic. When there is a failure, there is typically a usability problem. Studies of heuristics have shown them to be effective (in terms of numbers of problems found) and efficient (in terms of cost to perform) [48, 50].
Some researchers have illustrated difficulties with heuristic evaluation. Cockton & Woolrych suggest that heuristics should be used less in evaluation, in favor of empirical evaluations involving users [21]. Their arguments revolve around discrepancies among different evaluators and the low number of major problems that are found through the technique. Gray & Salzman also point out this weakness in [32].
Despite these objections, heuristic evaluation methods, particularly Nielsen’s, are still popular for their “discount” [70] approach to usability evaluation. Several recent works deal with adapting heuristic approaches to specified areas.
Baker et al. report on adapting heuristic evaluation to groupware systems [5]. They show that applying heuristic evaluation methods to groupware systems is effective and efficient for formative usability evaluation.
Mankoff et al. actually compare an adapted set of heuristics to Nielsen’s original set [56]. They studied ambient displays (which are similar to the systems that would be classified as ambient displays in the IRC framework) with both sets of heuristics and determined that their adapted set is better suited to ambient displays.
The heuristic usability evaluation method will be investigated in this research, but with different forms of heuristics, some adapted specifically to large screen information exhibits, others geared towards more general interface types (like generic notification systems or simply interfaces).
The focus of our work is to create a new set of heuristics by reliance on critical parameters. One that is tailored to the LSIE system class.
2.2.3 Comparing UEMs
Recent examples of work that strives to compare heuristic approaches to other UEMs (like lab-based user testing) include work shown at the 46th Annual Meeting of the Human Factors and Ergonomics Society.
Chattratichart and Brodie report on a comparison study of heuristic methods [16]. They extended heuristic evaluation (based on Nielsen’s) with a small set of content areas. These content areas served to focus the evaluation, thus producing more reliable results. It should also be noted that subjective opinions about the new method favored the original approach over the new approach. The added complexity of grouping problems into the content areas is the speculated cause of this finding [16].
Tan and Bishu compared heuristic evaluation to user testing [90]. They focused their work on web page evaluation and found that heuristic evaluation found more problems, but that the two techniques found different classes of problems. This means that these two methods are difficult to compare since the resulting problem lists are so different (like comparing apples to oranges). This difficulty in comparing analytical to empirical methods has been debated (see Human Computer Interaction 13(4) for a great summary of this debate) before and this particular work brings it to light in a more current example.
There has been some work on the best ways to compare UEMs. These studies are often limited to a specific area within HCI.
For example, Lavery et al. compared heuristics and task analysis in the domain of software visualization [52]. Their work resulted in development of problem reports that facilitate comparison of problems found with different methods. Their comparisons relied on effectiveness, efficiency, and validity measures for each method.
Others have also pointed out that effectiveness, efficiency, and validity are desirable measures for comparing UEMs (beyond simple numbers of usability problems obtained through the method) [40, 21].
Hartson et al. further put forth thoroughness, validity, reliability, and downstream utility as measures for comparing usability evaluation methods [40].
Chapter 3
Background and Motivation
However, there are many different types of usability evaluation methods one could employ to test design, and it is unclear which ones would serve as the best for this system class (large screen information exhibits).
One important variation in methods is whether to use an interface-specific tool or a generic tool that applies to a broad class of systems.
This preliminary study investigates tradeoffs of these two approaches (generic or specific) for evaluating LSIEs, by applying two types of evaluation to example LSIE systems.
This work provides the motivation and direction for the creation, testing, and use of a new set of heuristics tailored to the LSIE system class.
3.2 Assessing Evaluation Methods
Specific evaluation tools are developed for a single application, and apply solely to the system being tested (we refer to this as a per-study basis).
Many researchers use this approach, creating evaluation metrics, heuristics, or questionnaires tailored to the system in question (for example see [5, 56]). These tools seem advantageous because they provide fine grained insight into the target system, yielding detailed redesign solutions. However, filling immediate needs is costly—for each system to be tested a new evaluation method needs to be designed (by designers or evaluators), implemented, and used in the evaluation phase of software development.
In contrast, system-class evaluation tools are not tailored to a specific system and tend to focus on higher level, critical problem areas that might occur in systems within a common class.
These methods are created once (by usability experts) and used many times in separate evaluations. They are desirable for allowing ready application, promoting comparison between different systems, benchmarking system performance measures, and recognizing long-term, multi-project development progress.
However, using a system-class tool often means evaluators sacrifice focus on important interface details, since not all of the system aspects may be addressed by a generic tool. The appeal of system- class methods is apparent over a long-term period, namely through low cost and high benefit.
We conducted an experiment to determine the benefits of each approach in supporting a claims analysis, a key process within the scenario-based design approach [15, 77]. In a claims analysis, an evaluator makes claims about how important interface features will impact users.
Claims can be expressed as tradeoffs, conveying upsides or downsides of interface aspects like supported or unsupported activities, use of metaphors, information design choices (use of color, audio, icons, etc.), or interaction design techniques (affordances, feedback, configuration options, etc.). These claims capture the psychological impacts specific design decisions may have on users.
3.3 Motivation from Prior Work
UEM research efforts have developed high level, generic evaluation procedures, a notable example being Nielsen’s heuristics [70].
Heuristic evaluation has been embraced by practitioners because of its discount approach to assessing usability. With this approach (which involves identification of usability problems that fall into nine general and “most common problem areas”), 3-5 expert evaluators can uncover 70% of an interface’s usability problems.
However, the drawbacks to this approach (and most generic approaches) are evident in the need to develop more specific versions of heuristics for particular classes of systems.
For example, Mankoff et al. created a modified set of heuristics for ambient displays [56]. These displays differ from regular interfaces in that they often reside off the desktop, incorporating parts of the physical space in their design, hence necessitating a more specific approach to evaluation. They came up with the new set of heuristics by eliminating some from Nielsen’s original set, modifying the remaining heuristics to reflect ambient wording, and then added five new heuristics [56]. However, they do not report the criteria used in eliminating the original heuristics, the reasons for using the new wordings, or how they came up with the five new heuristics. They proceeded to compare this new set of heuristics to Nielsen’s original set and found the more specific heuristics provided better usability results.
Similar UEM work dealt with creating modified heuristics for groupware systems [5]. In this work, Baker et al. modified Nielsen’s original set to more closely match the user goals and needs associated with groupware systems. They based their modification on prior groupware system models to provide guidance in modifying Nielsen’s heuristics. The Locales Framework [35] and the mechanics of collaboration [38] helped Baker et al. in formulating their new heuristics. However, they do not describe how these models helped them in their creation, nor how they were used. From the comparison, they found the more application class-specific set of heuristics produced better results compared to the general set (Nielsen’s).
Both of these studies suggest that system-class specific heuristics are more desirable for formative evaluation. However, the creation processes used in both are not adequately described. It seems that to obtain the new set of heuristics, all the researchers did was modify Nielsen’s heuristics.
Unfortunately, it is not clear how this modification occurred. Did the researchers base the changes on important user goals for the system, as determined through critical parameters for the system class? Or was the modification based on guesswork or simple “this seems important for this type of system” style logic?
Source: Somervell, Jacob. Developing Heuristic Evaluation Methods for Large Screen Information Exhibits Based on Critical Parameters. [Dissertation, PhD in Computer Science and Applications] Virginia Polytechnic Institute and State University. June 22, 2004.
Subscribe to:
Posts (Atom)

