My comments: This is a good summary/intro of several Usability Questionnaires.
It has been observed that questionnaires are the most frequently used tools for usability evaluation. This page is a list of usability questionnaire resources, extending the information presented on the questionnaires page of Usabilitynet.
SUMI
This is a mature questionnaire whose standardisation base and manual have been regularly updated. It is good for desktop products, but has also been used to evaluate command-and-control applications. It is a commercial product which comes complete with scoring and report generation software. It is designed and sold by the Human Factors Research Group at University College Cork.
WAMMI
This is a new questionnaire, designed to evaluate the quality of use of web sites. It is backed up by an extensive standardisation database, and it is purchased on a per report basis. It is the result of a joint development project by Jurek Kirakowski and Nigel Claridge.
SUS
This is a mature questionnaire, developed by John Brooke in 1986 and not published until years later. It is very robust and has been extensively used and adapted. It is public domain and nobody has published any standardisation data about it. Of all the public domain questionnaires, this is the most strongly recommended.
QUIS
This is a questionnaire developed by Kent Norman that has been modified many times to keep it current since its first appearance. It is commercially available and is championed by Ben Shneiderman in his book Designing the User Interface. Reading, MA: Addison-Wesley Publishing Co., 1998. Despite lack of standardisation and validation data, it has many adherents.
USE (see http://www.mindspring.com/~alund/USE/IntroductionToUse.html)
This questionnaire is still in development by Arnie Lund (last updated 11/11/98). It attempts to create a three-factor model of usability that can be applied to many situations. However, no reliability or validation data are presented. Public domain use is encouraged.
CSUQ
This is a well-designed questionnaire developed by Jim Lewis and it is public domain. It has excellent psychometric reliability properties but no standardisation base.
IsoNorm (in German only)
This questionnaire is designed to test the usability quality of software following the ISO 9241 part 10 principles. It is created by a team led by Jochim Puemper. Strong reliabilities are claimed for the sub-scales, although it appears there may be a strong inter-correlation between them as well. Downloads and an on-line version are available from the above URL, as well as articles about it (all in German.)
IsoMetrics
This questionnaire is produced by Guenter Gediga and his team. It is another attempt to produce a way of measuring ISO 9241 part 10, with reference to specific software features that may give rise to low usability data. It is therefore good both for summative and formative assessments. The questionnaire is well researched and detailed statistical information is given. Downloads of English and German versions are available. There is no standardisation base for it but it is public domain.
Questionnaire resources
Source: http://www.usabilitynet.org/tools/r_questionnaire.htm
Showing posts with label SUMI. Show all posts
Showing posts with label SUMI. Show all posts
Friday, October 9, 2009
Thursday, October 8, 2009
Oct 9 - HCIRN, Usability Questionnaires
Usability questionnaires intend to assess users' perception of the usability of a product.
A number of usability questionnaires have been developed or are currently under development:
ASQ (After-Scenario Questionnaire)
CSUQ (Computer System Usability Questionnaire)
CUSI (Computer User Satisfaction Inventory)
ErgoNorm Questionnaire
IsoMetrics
ISONORM 9241/10
MUMMS (Measuring the Usability of Multi-Media Systems)
PSSUQ (Post-Study System Usability Questionnaire)
PUEU (Perceived Usefulness and Ease of Use)
PUTQ (Purdue Usability Testing Questionnaire)
QUIS (Questionnaire for User Interaction Satisfaction)
SUMI (Software Usability Measurement Inventory)
SUS (System Usability Scale)
USE (Usefulness, Satisfaction, and Ease of Use)
WAMMI (Website Analysis and MeasureMent Inventory)
Some of these questionnaires are free, others are commercial and require a license to use them.
Two popular commercial questionnaires are QUIS (Questionnaire for User Interaction Satisfaction) and SUMI (Software Usability Measurement Inventory). Both provided analysis software with output in various formats.
Two popular free questionnaires are PSSUQ (Post-Study System Usability Questionnaire) and SUS (System Usability Scale).
Two of the questionnaires in the above list are related to the seven dialog principles described in ISO 9241 Part 10. IsoMetrics is directly based on these dialog principles and intends to measure the usability of a system along these dimensions. SUMI has five factors, at least four of which seem to directly correspond to dialog principles.
Global vs. Analytic Evaluation
Usability questionnaires generally provide only global measures of system usability. They do not provide any information about which aspects of the system contribute to the positive or negative evaluation. (Many authors use the terms formative and summative evaluation to distinguish between these types evaluation. However, based on a detailed analysis of these terms we believe that the terms analytic and global evaluation are more appropriate. See Formative and Summative Evaluation for more detail.)
Providing only a global measure is a major limitation of usability questionnaires. A number of questionnaires, therefore, attempt to facilitate analytic evaluation as well. One common way which nearly all usability questionnaires support is to provide a comments sections in which respondents can enter feedback in free form.
Some questionnaires also support some form of item analysis. A number of case studies using QUIS (Questionnaire for User Interaction Satisfaction) compared average item scores against the overall average score. However, this approach is questionable.
SUMI (Software Usability Measurement Inventory) uses a method called Item Consensual Analysis (ICA) to obtain additional information about usability problems. Item Consensual Analysis compares the response pattern of each item with its response pattern in the standardization database.
Assessing a Single System vs. Comparing Two or More Systems
Questionnaires can be used to assess a single system or to compare two or more systems. When using a questionnaire to assess a single system, the resulting score by itself is not very meaningful. What does it mean that your system has a score of 5.0 on a range from 1.0 to 7.0? It is better or worse than other comparable systems? To interpret a single score, you need a baseline to compare it with.
However, for a baseline to be useful, it has to be updated regularly. If the baseline score was 5.0 two years ago, it could be quite different now. Improvements to applications could have increased the baseline score to 6.0. Conversely, changes to applications may have increased their complexity which decreased the baseline score to 4.0.
Also, baselines need to be specific for the type of application. For example, new products, such as web browsers, may initially have a poorer score than established products, such as word processors. So a score of 5.0 might be poor compared to word processors and excellent compared to web browsers.
Finally, different baselines are needed for different countries. There might be cultural differences in the interpretation of specific questions as well as in the response tendency of the respondents. This makes it impossible to compare a score to a baseline obtained for another language version of the same questionnaire.
To our knowledge, to date only SUMI (Software Usability Measurement Inventory) provides such a standardization database.
Baselines are not required when comparing two or more systems. All usability questionnaires, provided they have sufficient reliability, are suitable for this. They can be used to measure changes in usability between versions of the same system or to assess differences between competing products.
Validity of Usability Questionnaires
Usability questionnaires can be interpreted in two ways. Firstly, they are a direct measure of users' perception of the usability of a system. Secondly, they are an indirect measure of the actual usability of a system.
...............
...............
One way to validate a questionnaire is to compare it to other questionnaires that claim to measure the same construct. However, few such studies have been conducted for usability questionnaires. Kirakowski (1994) reports two studies that directly compared CUSI (Computer User Satisfaction Inventory), SUS (System Usability Scale) and QUIS (Questionnaire for User Interaction Satisfaction) 5.0. Lucey (1991) found good correlations between the affect subscale of CUSI with SUS and overall QUIS. The correlations between the competence subscale of CUSI with SUS and QUIS were low. Wong & Rengger (1990) found a similar pattern.
Choosing a Usability Questionnaire
Which usability questionnaire should you choose? We have not had the opportunity to examine SUMI (Software Usability Measurement Inventory).
SUMI has never been published and is only available by purchasing the questionnaire package. However, it looks to be your best choice. It is a well-designed and extensively tested questionnaire. It is available in various languages and is the only usability questionnaire that has been standardized. Besides producing scores for a global scale and five subscales, detailed information about a system can be obtained by analyzing the response pattern of each item. The only disadvantage of SUMI is that it costs money. It is worth it, but HCI practitioners might find it difficult to convince management that the expense is justified.
Out of the freely available questionnaires, SUS (System Usability Scale) is probably your best choice. It does not have the standardization database of SUMI and it provides only an overall usability score. However, it is short and less susceptible to response bias than some other freely available questionnaires. SUS is an excellent instrument for comparing systems.
Compared to other evaluation methods in HCI, there is little literature available on usability questionnaires. There are also a number of weak spots in the published research on usability questionnaires. In particular, extensive validation studies that compare usability questionnaires with each other and with other usability evaluation methods are missing. The reason for this is not clear. Perhaps this is an indication that usability questionnaires are currently of little relevance in HCI practice.
Kirakowski, 1994 contains a good introduction to usability questionnaires.
Links
Questionnaires in Usability Engineering: A List of Frequently Asked Questions
By Jurek Kirakowski
http://www.ucc.ie/hfrg/resources/qfaq1.html
UsabilityNet > Questionnaire Resources
http://www.usabilitynet.org/tools/r_questionnaire.htm
A short overview of usability questionnaires.
Web-Based User Interface Evaluation with Questionnaires
By Gary Perlman
http://www.acm.org/~perlman/question.html
Includes online versions of a number of usability questionnaires.
Source: http://www.hcirn.com/atoz/atozu/usaques.php
A number of usability questionnaires have been developed or are currently under development:
ASQ (After-Scenario Questionnaire)
CSUQ (Computer System Usability Questionnaire)
CUSI (Computer User Satisfaction Inventory)
ErgoNorm Questionnaire
IsoMetrics
ISONORM 9241/10
MUMMS (Measuring the Usability of Multi-Media Systems)
PSSUQ (Post-Study System Usability Questionnaire)
PUEU (Perceived Usefulness and Ease of Use)
PUTQ (Purdue Usability Testing Questionnaire)
QUIS (Questionnaire for User Interaction Satisfaction)
SUMI (Software Usability Measurement Inventory)
SUS (System Usability Scale)
USE (Usefulness, Satisfaction, and Ease of Use)
WAMMI (Website Analysis and MeasureMent Inventory)
Some of these questionnaires are free, others are commercial and require a license to use them.
Two popular commercial questionnaires are QUIS (Questionnaire for User Interaction Satisfaction) and SUMI (Software Usability Measurement Inventory). Both provided analysis software with output in various formats.
Two popular free questionnaires are PSSUQ (Post-Study System Usability Questionnaire) and SUS (System Usability Scale).
Two of the questionnaires in the above list are related to the seven dialog principles described in ISO 9241 Part 10. IsoMetrics is directly based on these dialog principles and intends to measure the usability of a system along these dimensions. SUMI has five factors, at least four of which seem to directly correspond to dialog principles.
Global vs. Analytic Evaluation
Usability questionnaires generally provide only global measures of system usability. They do not provide any information about which aspects of the system contribute to the positive or negative evaluation. (Many authors use the terms formative and summative evaluation to distinguish between these types evaluation. However, based on a detailed analysis of these terms we believe that the terms analytic and global evaluation are more appropriate. See Formative and Summative Evaluation for more detail.)
Providing only a global measure is a major limitation of usability questionnaires. A number of questionnaires, therefore, attempt to facilitate analytic evaluation as well. One common way which nearly all usability questionnaires support is to provide a comments sections in which respondents can enter feedback in free form.
Some questionnaires also support some form of item analysis. A number of case studies using QUIS (Questionnaire for User Interaction Satisfaction) compared average item scores against the overall average score. However, this approach is questionable.
SUMI (Software Usability Measurement Inventory) uses a method called Item Consensual Analysis (ICA) to obtain additional information about usability problems. Item Consensual Analysis compares the response pattern of each item with its response pattern in the standardization database.
Assessing a Single System vs. Comparing Two or More Systems
Questionnaires can be used to assess a single system or to compare two or more systems. When using a questionnaire to assess a single system, the resulting score by itself is not very meaningful. What does it mean that your system has a score of 5.0 on a range from 1.0 to 7.0? It is better or worse than other comparable systems? To interpret a single score, you need a baseline to compare it with.
However, for a baseline to be useful, it has to be updated regularly. If the baseline score was 5.0 two years ago, it could be quite different now. Improvements to applications could have increased the baseline score to 6.0. Conversely, changes to applications may have increased their complexity which decreased the baseline score to 4.0.
Also, baselines need to be specific for the type of application. For example, new products, such as web browsers, may initially have a poorer score than established products, such as word processors. So a score of 5.0 might be poor compared to word processors and excellent compared to web browsers.
Finally, different baselines are needed for different countries. There might be cultural differences in the interpretation of specific questions as well as in the response tendency of the respondents. This makes it impossible to compare a score to a baseline obtained for another language version of the same questionnaire.
To our knowledge, to date only SUMI (Software Usability Measurement Inventory) provides such a standardization database.
Baselines are not required when comparing two or more systems. All usability questionnaires, provided they have sufficient reliability, are suitable for this. They can be used to measure changes in usability between versions of the same system or to assess differences between competing products.
Validity of Usability Questionnaires
Usability questionnaires can be interpreted in two ways. Firstly, they are a direct measure of users' perception of the usability of a system. Secondly, they are an indirect measure of the actual usability of a system.
...............
...............
One way to validate a questionnaire is to compare it to other questionnaires that claim to measure the same construct. However, few such studies have been conducted for usability questionnaires. Kirakowski (1994) reports two studies that directly compared CUSI (Computer User Satisfaction Inventory), SUS (System Usability Scale) and QUIS (Questionnaire for User Interaction Satisfaction) 5.0. Lucey (1991) found good correlations between the affect subscale of CUSI with SUS and overall QUIS. The correlations between the competence subscale of CUSI with SUS and QUIS were low. Wong & Rengger (1990) found a similar pattern.
Choosing a Usability Questionnaire
Which usability questionnaire should you choose? We have not had the opportunity to examine SUMI (Software Usability Measurement Inventory).
SUMI has never been published and is only available by purchasing the questionnaire package. However, it looks to be your best choice. It is a well-designed and extensively tested questionnaire. It is available in various languages and is the only usability questionnaire that has been standardized. Besides producing scores for a global scale and five subscales, detailed information about a system can be obtained by analyzing the response pattern of each item. The only disadvantage of SUMI is that it costs money. It is worth it, but HCI practitioners might find it difficult to convince management that the expense is justified.
Out of the freely available questionnaires, SUS (System Usability Scale) is probably your best choice. It does not have the standardization database of SUMI and it provides only an overall usability score. However, it is short and less susceptible to response bias than some other freely available questionnaires. SUS is an excellent instrument for comparing systems.
Compared to other evaluation methods in HCI, there is little literature available on usability questionnaires. There are also a number of weak spots in the published research on usability questionnaires. In particular, extensive validation studies that compare usability questionnaires with each other and with other usability evaluation methods are missing. The reason for this is not clear. Perhaps this is an indication that usability questionnaires are currently of little relevance in HCI practice.
Kirakowski, 1994 contains a good introduction to usability questionnaires.
Links
Questionnaires in Usability Engineering: A List of Frequently Asked Questions
By Jurek Kirakowski
http://www.ucc.ie/hfrg/resources/qfaq1.html
UsabilityNet > Questionnaire Resources
http://www.usabilitynet.org/tools/r_questionnaire.htm
A short overview of usability questionnaires.
Web-Based User Interface Evaluation with Questionnaires
By Gary Perlman
http://www.acm.org/~perlman/question.html
Includes online versions of a number of usability questionnaires.
Source: http://www.hcirn.com/atoz/atozu/usaques.php
Oct 9 - Fifty statements of SUMI
SUMI - Software Usability Measurement Inventory
1 This software responds too slowly to inputs.
2 I would recommend this software to my colleagues.
3 The instructions and prompts are helpful.
4 The software has at some time stopped unexpectedly.
5 Learning to operate this software initially is full of problems.
6 I sometimes don't know what to do next with this software.
7 I enjoy my sessions with this software.
8 I find that the help information given by this software is not very useful.
9 If this software stops, it is not easy to restart it.
10 It takes too long to learn the software commands.
11 I sometimes wonder if I'm using the right command.
12 Working with this software is satisfying.
13 The way that system information is presented is clear and understandable.
14 I feel safer if I use only a few familiar commands or operations.
15 The software documentation is very informative.
16 This software seems to disrupt the way I normally like to arrange my work.
17 Working with this software is mentally stimulating.
18 There is never enough information on the screen when it’s needed.
19 I feel in command of this software when I am using it.
20 I prefer to stick to the facilities that I know best.
21 I think this software is inconsistent.
22 I would not like to use this software every day.
23 I can understand and act on the information provided by this software.
24 This software is awkward when I want to do something which is not standard.
25 There is too much to read before you can use the software.
26 Tasks can be performed in a straightforward manner using this software.
27 Using this software is frustrating.
28 The software has helped me overcome any problems I have had in using it.
29 The speed of this software is fast enough.
30 I keep having to go back to look at the guides.
31 It is obvious that user needs have been fully taken into consideration.
32 There have been times in using this software when I have felt quite tense.
33 The organisation of the menus or information lists seems quite logical.
34 The software allows the user to be economic of keystrokes.
35 Learning how to use new functions is difficult.
36 There are too many steps required to get something to work.
37 I think this software has made me have a headache on occasion.
38 Error prevention messages are not adequate.
39 It is easy to make the software do exactly what you want.
40 I will never learn to use all that is offered in this software.
41 The software hasn’t always done what I was expecting.
42 The software has a very attractive presentation.
43 Either the amount or quality of the help information varies across the system.
44 It is relatively easy to move from one part of a task to another.
45 It is easy to forget how to do things with this software.
46 This software occasionally behaves in a way which can’t be understood.
47 This software is really very awkward.
48 It is easy to see at a glance what the options are at each stage.
49 Getting data files in and out of the system is not easy.
50 I have to look for assistance most times when I use this software.
Five categories are:
* Efficiency
* Affect
* Helpfulness
* Control
* Learnability.
Each category has 10 statements, as could be seen above.
Source: http://sumi.ucc.ie/uksample.pdf
1 This software responds too slowly to inputs.
2 I would recommend this software to my colleagues.
3 The instructions and prompts are helpful.
4 The software has at some time stopped unexpectedly.
5 Learning to operate this software initially is full of problems.
6 I sometimes don't know what to do next with this software.
7 I enjoy my sessions with this software.
8 I find that the help information given by this software is not very useful.
9 If this software stops, it is not easy to restart it.
10 It takes too long to learn the software commands.
11 I sometimes wonder if I'm using the right command.
12 Working with this software is satisfying.
13 The way that system information is presented is clear and understandable.
14 I feel safer if I use only a few familiar commands or operations.
15 The software documentation is very informative.
16 This software seems to disrupt the way I normally like to arrange my work.
17 Working with this software is mentally stimulating.
18 There is never enough information on the screen when it’s needed.
19 I feel in command of this software when I am using it.
20 I prefer to stick to the facilities that I know best.
21 I think this software is inconsistent.
22 I would not like to use this software every day.
23 I can understand and act on the information provided by this software.
24 This software is awkward when I want to do something which is not standard.
25 There is too much to read before you can use the software.
26 Tasks can be performed in a straightforward manner using this software.
27 Using this software is frustrating.
28 The software has helped me overcome any problems I have had in using it.
29 The speed of this software is fast enough.
30 I keep having to go back to look at the guides.
31 It is obvious that user needs have been fully taken into consideration.
32 There have been times in using this software when I have felt quite tense.
33 The organisation of the menus or information lists seems quite logical.
34 The software allows the user to be economic of keystrokes.
35 Learning how to use new functions is difficult.
36 There are too many steps required to get something to work.
37 I think this software has made me have a headache on occasion.
38 Error prevention messages are not adequate.
39 It is easy to make the software do exactly what you want.
40 I will never learn to use all that is offered in this software.
41 The software hasn’t always done what I was expecting.
42 The software has a very attractive presentation.
43 Either the amount or quality of the help information varies across the system.
44 It is relatively easy to move from one part of a task to another.
45 It is easy to forget how to do things with this software.
46 This software occasionally behaves in a way which can’t be understood.
47 This software is really very awkward.
48 It is easy to see at a glance what the options are at each stage.
49 Getting data files in and out of the system is not easy.
50 I have to look for assistance most times when I use this software.
Five categories are:
* Efficiency
* Affect
* Helpfulness
* Control
* Learnability.
Each category has 10 statements, as could be seen above.
Source: http://sumi.ucc.ie/uksample.pdf
Oct 9 - SUMI: Software Usability Measurement Inventory
SUMI - Software Usability Measurement Inventory
The de facto industry standard evaluation questionnaire for assessing quality of use of software by end users.
What is SUMI
The Software Usability Measurement Inventory is a rigorously tested and proven method of measuring software quality from the end user's point of view.
SUMI is a consistent method for assessing the quality of use of a software product or prototype, and can assist with the detection of usability flaws before a product is shipped.
It is backed by an extensive reference database embedded in an effective analysis and report generation tool.
Who should use SUMI?
SUMI is recommended to any organisation which wishes to measure the perceived quality of use of software, either as a developer, a consumer of software, or as a purchaser/consultant. SUMI is increasingly being used to set quality of use requirements by software procurers.
SUMI also assists the manager in identifying the most appropriate software for their organisation. It has been well documented that if staff have quality tools to work with, this contributes to overall efficiency of staff and the quality of their work output.
Our customers have used SUMI effectively to:
* assess new products during product evaluation
* make comparisons between products or versions of products
* set targets for future application developments.
SUMI has been used specifically within development environments to:
* set verifiable goals for quality of use attainment
* track achievement of targets during product development
* highlight good and bad aspects of an interface.
SUMI is the de facto industry standard questionnaire for analysing users' responses to desktop software or software applications provided through the internet.
Why use SUMI?
SUMI is the only commercially available questionnaire for the assessment of the usability of software which has been developed, validated, and standardised on an international basis.
There is a large range of languages in which SUMI is available. Each language version is carefully translated and validated by native speakers of the language.
SUMI ennables measurement of some of the user-orientated requirements expressed in the European Directive on Minimum Health and Safety Requirements for Work with Display Screen Equipment (90/270/EEC).
SUMI is mentioned in the ISO 9241 standard as a recognised method of testing user satisfaction.
What does SUMI look like?
SUMI consists of 50 statements to which the user has to reply that they either Agree, Don't Know, or Disagree.
Here are some example statements: Item No. Item Wording
1. This software responds too slowly to inputs.
3. The instructions and prompts are helpful.
13. The way that system information is presented is clear and understandable.
22. I would not like to use this software every day.
You may also take a look at a sample questionnaire (UK wording) in pdf format,
How do I administer SUMI and how long does it take?
It takes a user about 3 minutes to fill out the questionnaire, perhaps a few minutes longer on the internet version.
One way of administering it is on paper: print out the SUMI form and get your user to make marks on the page. It takes an analyst about one minute to type each user's responses into a file for scoring by SUMISCO or to send to HFRG for scoring.
On the other hand you may decide to go for the internet, online option.
You can also do a hybrid: concoct your own HTML pages to be served on your intranet and either send the results to HFRG for analysis, or purchase SUMISCO and analyse them yourself.
How many users do I need?
Online SUMI might require sample sizes with a minimum of about 30 unless your respondents are well targetted.
However, we know that paper SUMI will give you reliable results with as few as 12 users. This is because you are able to control the quality of your user sample directly when administering SUMI on paper.
You can use fewer users if you wish, but beware that your results may not be as representative of the true user population. In fact, SUMI has yielded useful information with user sample sizes of four or five.
However, this question is always like 'how long is a piece of string.' You should try to get as many users as you can within your timeframe.
Source: http://sumi.ucc.ie/whatis.html
3. SUMI Development
Work on SUMI started in late 1990. One of the work packages entrusted to the HFRG within the MUSiC project was to develop questionnaire methods of assessing usability. The objectives of this work package were:
* to examine the CUSI Competence scale and to expand it and to extract further subscales if warranted by the evidence;
* to achieve an international standardisation database for the new questionnaire and to validate its use in commercial environments.
Both these objectives were achieved by the end of the project. The SUMI questionnaire was first published in 1993 and has been widely disseminated since then, both in Europe and in the United States.
3.1 Psychometric development
SUMI started with an initial item pool of over 150 items, assembled from previously reported studies (including many reviewed above), from discussions with actual end users about their experiences with information technology, and from suggestions given by experts in HCI and software engineers working in the MUSiC project. The items were examined for consistency of perceived meaning by getting 10 subject matter experts to allocate each item to content areas. Items were then rewritten or eliminated if they produced inconsistent allocations.
The questionnaire developers opted for a Lickert scaling approach, both for historical reasons (the CUSI questionnaire was Lickert-scaled) and because this is considered to be a natural way of eliciting opinions about a software product. Different types of scales in use in questionnaire design within HCI are discussed in Kirakowski and Corbett (1990). The implication is that each item is considered to have roughly similar importance, and that the strength of a user's opinion can be estimated by summing or averaging the individual ratings of strength of opinion for each item. Many items are used in order to overcome variability due to extraneous or irrelevant factors.
This procedure produced the first questionnaire form, which consisted of 75 satisfactory items. The respondents had to decide whether they agreed strongly, agreed, didn't know, disagreed, or disagreed strongly with each of the 75 items in relation to the software they were evaluating.
Questionnaires were administered to 139 end users from a range of organisations (this was called sample 1). The respondents completed the inventory at their work place with the software they were evaluating near at hand. All these respondents were genuine end users who were using the software to accomplish task goals within their organisations for their daily work. The resulting matrix of inter-correlations between items was factor analysed and the items were observed to relate to a number of different meaningful areas of user perception of usability. Five to six groupings emerged which gave acceptable measures of internal consistency and score distributions.
Revisions were made to some items to centralise means and improve item variances, and then the ten best items with highest factor loadings were retained for each grouping. The number of groups of items was set to five. Items were revised in the light of critique from the industrial partners of MUSiC in order to reflect the growing trend towards Graphical User Interfaces. A number of users had remarked that it was difficult to make a judgement over five categories of response for some items. After some discussion, it was decided to change the response categories to three: Agree, Don't Know, and Disagree.
This produced the second questionnaire form of 50 items, in which each subscale was represented by 10 different items.
Typical items from this version are:Item No. Item Wording
1. This software responds too slowly to inputs.
3. The instructions and prompts are helpful.
13. The way that system information is presented is clear and understandable.
22. I would not like to use this software every day.
A new sample (sample 2) of data from 143 users in a commercial environment was collected. Analysis of this sample of data showed that item response rates, scale reliabilities, and item-scale correlations were similar to or better than those in the first form's sample. Analyses of variance showed that the questionnaire differentiated between different software systems in the sample. After analysis of sample 2, a few items were revised slightly to improve their scale properties. The subscales were given descriptive labels by the questionnaire developers. These were:
* Efficiency
*Affect
* Helpfulness
* Control
* Learnability.
The precise meaning of these subscales is given in the SUMI manual, but in general, the Affect subscale measures (as before, in CUSI) the user's general emotional reaction to the software -- it may be glossed as Likeability. Efficiency measures the degree to which users feel that the software assists them in their work and is related to the concept of transparency. Helpfulness measures the degree to which the software is self-explanatory, as well as more specific things like the adequacy of help facilities and documentation. The Control dimensions measures the extent to which the user feels in control of the software, as opposed to being controlled by the software, when carrying out the task. Learnability, finally, measures the speed and facility with which the user feels that they have been able to master the system, or to learn how to use new features when necessary.
At this time, a validity study was carried out in a company who has requested to remain anonymous. In this company, two versions of the same editing software were installed for the use of the programmer teams. The users carried out the same kinds of tasks in the same environments, but each programmer team used only one of software versions for most of the time. There were 20 users for version 1, and 22 for version 2. Analysis of variance showed a significant effect for the SUMI scales, for the difference between systems, and for the interaction between SUMI scales and systems. This last finding was important as it indicated that the SUMI scales were not responding en masse but that they were discriminating between differential levels of components of usability. Table 1 shows the means and standard deviations of the SUMI scales and the two software versions.
In fact, Version 2 was considerably more popular among its users, and the users of this version considered that they were able to carry out their tasks more efficiently with it. Learnability was not considered to be an issue by either group of users. The Data Processing manager of the company reviewed the results with the questionnaire development team and in his opinion, the results confirmed informal feedback and observation. Later we learnt that the company subsequently decided to switch to Version 2, not only on the basis of our results, but because that was the general feeling.
At this stage, the Global scale was also derived. The Global scale consists of 25 items out of the 50 which loaded most heavily on a general usability factor. Because of the larger number of items which contribute to the Global scale, reliabilities are correspondingly higher. The Global scale was produced in order to represent the single construct of perceived quality of use better than a simple average of all the items of the questionnaire.
...................
...................
Source: http://sumi.ucc.ie/sumipapp.html
The de facto industry standard evaluation questionnaire for assessing quality of use of software by end users.
What is SUMI
The Software Usability Measurement Inventory is a rigorously tested and proven method of measuring software quality from the end user's point of view.
SUMI is a consistent method for assessing the quality of use of a software product or prototype, and can assist with the detection of usability flaws before a product is shipped.
It is backed by an extensive reference database embedded in an effective analysis and report generation tool.
Who should use SUMI?
SUMI is recommended to any organisation which wishes to measure the perceived quality of use of software, either as a developer, a consumer of software, or as a purchaser/consultant. SUMI is increasingly being used to set quality of use requirements by software procurers.
SUMI also assists the manager in identifying the most appropriate software for their organisation. It has been well documented that if staff have quality tools to work with, this contributes to overall efficiency of staff and the quality of their work output.
Our customers have used SUMI effectively to:
* assess new products during product evaluation
* make comparisons between products or versions of products
* set targets for future application developments.
SUMI has been used specifically within development environments to:
* set verifiable goals for quality of use attainment
* track achievement of targets during product development
* highlight good and bad aspects of an interface.
SUMI is the de facto industry standard questionnaire for analysing users' responses to desktop software or software applications provided through the internet.
Why use SUMI?
SUMI is the only commercially available questionnaire for the assessment of the usability of software which has been developed, validated, and standardised on an international basis.
There is a large range of languages in which SUMI is available. Each language version is carefully translated and validated by native speakers of the language.
SUMI ennables measurement of some of the user-orientated requirements expressed in the European Directive on Minimum Health and Safety Requirements for Work with Display Screen Equipment (90/270/EEC).
SUMI is mentioned in the ISO 9241 standard as a recognised method of testing user satisfaction.
What does SUMI look like?
SUMI consists of 50 statements to which the user has to reply that they either Agree, Don't Know, or Disagree.
Here are some example statements: Item No. Item Wording
1. This software responds too slowly to inputs.
3. The instructions and prompts are helpful.
13. The way that system information is presented is clear and understandable.
22. I would not like to use this software every day.
You may also take a look at a sample questionnaire (UK wording) in pdf format,
How do I administer SUMI and how long does it take?
It takes a user about 3 minutes to fill out the questionnaire, perhaps a few minutes longer on the internet version.
One way of administering it is on paper: print out the SUMI form and get your user to make marks on the page. It takes an analyst about one minute to type each user's responses into a file for scoring by SUMISCO or to send to HFRG for scoring.
On the other hand you may decide to go for the internet, online option.
You can also do a hybrid: concoct your own HTML pages to be served on your intranet and either send the results to HFRG for analysis, or purchase SUMISCO and analyse them yourself.
How many users do I need?
Online SUMI might require sample sizes with a minimum of about 30 unless your respondents are well targetted.
However, we know that paper SUMI will give you reliable results with as few as 12 users. This is because you are able to control the quality of your user sample directly when administering SUMI on paper.
You can use fewer users if you wish, but beware that your results may not be as representative of the true user population. In fact, SUMI has yielded useful information with user sample sizes of four or five.
However, this question is always like 'how long is a piece of string.' You should try to get as many users as you can within your timeframe.
Source: http://sumi.ucc.ie/whatis.html
3. SUMI Development
Work on SUMI started in late 1990. One of the work packages entrusted to the HFRG within the MUSiC project was to develop questionnaire methods of assessing usability. The objectives of this work package were:
* to examine the CUSI Competence scale and to expand it and to extract further subscales if warranted by the evidence;
* to achieve an international standardisation database for the new questionnaire and to validate its use in commercial environments.
Both these objectives were achieved by the end of the project. The SUMI questionnaire was first published in 1993 and has been widely disseminated since then, both in Europe and in the United States.
3.1 Psychometric development
SUMI started with an initial item pool of over 150 items, assembled from previously reported studies (including many reviewed above), from discussions with actual end users about their experiences with information technology, and from suggestions given by experts in HCI and software engineers working in the MUSiC project. The items were examined for consistency of perceived meaning by getting 10 subject matter experts to allocate each item to content areas. Items were then rewritten or eliminated if they produced inconsistent allocations.
The questionnaire developers opted for a Lickert scaling approach, both for historical reasons (the CUSI questionnaire was Lickert-scaled) and because this is considered to be a natural way of eliciting opinions about a software product. Different types of scales in use in questionnaire design within HCI are discussed in Kirakowski and Corbett (1990). The implication is that each item is considered to have roughly similar importance, and that the strength of a user's opinion can be estimated by summing or averaging the individual ratings of strength of opinion for each item. Many items are used in order to overcome variability due to extraneous or irrelevant factors.
This procedure produced the first questionnaire form, which consisted of 75 satisfactory items. The respondents had to decide whether they agreed strongly, agreed, didn't know, disagreed, or disagreed strongly with each of the 75 items in relation to the software they were evaluating.
Questionnaires were administered to 139 end users from a range of organisations (this was called sample 1). The respondents completed the inventory at their work place with the software they were evaluating near at hand. All these respondents were genuine end users who were using the software to accomplish task goals within their organisations for their daily work. The resulting matrix of inter-correlations between items was factor analysed and the items were observed to relate to a number of different meaningful areas of user perception of usability. Five to six groupings emerged which gave acceptable measures of internal consistency and score distributions.
Revisions were made to some items to centralise means and improve item variances, and then the ten best items with highest factor loadings were retained for each grouping. The number of groups of items was set to five. Items were revised in the light of critique from the industrial partners of MUSiC in order to reflect the growing trend towards Graphical User Interfaces. A number of users had remarked that it was difficult to make a judgement over five categories of response for some items. After some discussion, it was decided to change the response categories to three: Agree, Don't Know, and Disagree.
This produced the second questionnaire form of 50 items, in which each subscale was represented by 10 different items.
Typical items from this version are:Item No. Item Wording
1. This software responds too slowly to inputs.
3. The instructions and prompts are helpful.
13. The way that system information is presented is clear and understandable.
22. I would not like to use this software every day.
A new sample (sample 2) of data from 143 users in a commercial environment was collected. Analysis of this sample of data showed that item response rates, scale reliabilities, and item-scale correlations were similar to or better than those in the first form's sample. Analyses of variance showed that the questionnaire differentiated between different software systems in the sample. After analysis of sample 2, a few items were revised slightly to improve their scale properties. The subscales were given descriptive labels by the questionnaire developers. These were:
* Efficiency
*Affect
* Helpfulness
* Control
* Learnability.
The precise meaning of these subscales is given in the SUMI manual, but in general, the Affect subscale measures (as before, in CUSI) the user's general emotional reaction to the software -- it may be glossed as Likeability. Efficiency measures the degree to which users feel that the software assists them in their work and is related to the concept of transparency. Helpfulness measures the degree to which the software is self-explanatory, as well as more specific things like the adequacy of help facilities and documentation. The Control dimensions measures the extent to which the user feels in control of the software, as opposed to being controlled by the software, when carrying out the task. Learnability, finally, measures the speed and facility with which the user feels that they have been able to master the system, or to learn how to use new features when necessary.
At this time, a validity study was carried out in a company who has requested to remain anonymous. In this company, two versions of the same editing software were installed for the use of the programmer teams. The users carried out the same kinds of tasks in the same environments, but each programmer team used only one of software versions for most of the time. There were 20 users for version 1, and 22 for version 2. Analysis of variance showed a significant effect for the SUMI scales, for the difference between systems, and for the interaction between SUMI scales and systems. This last finding was important as it indicated that the SUMI scales were not responding en masse but that they were discriminating between differential levels of components of usability. Table 1 shows the means and standard deviations of the SUMI scales and the two software versions.
In fact, Version 2 was considerably more popular among its users, and the users of this version considered that they were able to carry out their tasks more efficiently with it. Learnability was not considered to be an issue by either group of users. The Data Processing manager of the company reviewed the results with the questionnaire development team and in his opinion, the results confirmed informal feedback and observation. Later we learnt that the company subsequently decided to switch to Version 2, not only on the basis of our results, but because that was the general feeling.
At this stage, the Global scale was also derived. The Global scale consists of 25 items out of the 50 which loaded most heavily on a general usability factor. Because of the larger number of items which contribute to the Global scale, reliabilities are correspondingly higher. The Global scale was produced in order to represent the single construct of perceived quality of use better than a simple average of all the items of the questionnaire.
...................
...................
Source: http://sumi.ucc.ie/sumipapp.html
Oct 8 - SUMI, MUMMS, WAMMI
Human Factors Research Group (HFRG) has 3 usability questionnaires. The most popular is SUMI.
Software Usability Measurement Inventory (SUMI)
The SUMI is a rigorously tested and proven method of measuring software quality from the end user's point of view. SUMI is a consistent method for assessing the quality of use of a software product or prototype, and can assist with the detection of usability flaws before a product is shipped. It is backed by an extensive reference database embedded in an effective analysis and report generation tool.
Measuring the Usability of Multi-Media Systems (MUMMS)
This questionnaire is designed for evaluating quality of use of multi-media software products. MUMMS is still under development, and we welcome participants who would like to contribute to the development effort on a data provider's agreement - see inside!
Website Analysis and MeasureMent Inventory (WAMMI)
A short but very reliable questionnaire that tells you what your visitors think about your web site.
Source: http://www.ucc.ie/hfrg/questionnaires/
Software Usability Measurement Inventory (SUMI)
The SUMI is a rigorously tested and proven method of measuring software quality from the end user's point of view. SUMI is a consistent method for assessing the quality of use of a software product or prototype, and can assist with the detection of usability flaws before a product is shipped. It is backed by an extensive reference database embedded in an effective analysis and report generation tool.
Measuring the Usability of Multi-Media Systems (MUMMS)
This questionnaire is designed for evaluating quality of use of multi-media software products. MUMMS is still under development, and we welcome participants who would like to contribute to the development effort on a data provider's agreement - see inside!
Website Analysis and MeasureMent Inventory (WAMMI)
A short but very reliable questionnaire that tells you what your visitors think about your web site.
Source: http://www.ucc.ie/hfrg/questionnaires/
Oct 8 - Usability Questionnaires >>Perlman, User Interface Usability Evaluation with Web-Based Questionnaires
Usability Questionnaires
User Interface Usability Evaluation with Web-Based Questionnaires
Last updated: 2009-07-15 by Gary Perlman (hcibib.org/perlman/ perlman@acm.org)
Acronym > Instrument > Reference > Institution > Example
QUIS > Questionnaire for User Interface Satisfaction > Chin et al, 1988 > Maryland >27 questions
PUEU > Perceived Usefulness and Ease of Use > Davis, 1989 > IBM > 12 questions
NAU > Nielsen's Attributes of Usability > Nielsen, 1993 > Bellcore > 5 attributes
NHE > Nielsen's Heuristic Evaluation > Nielsen, 1993 > Bellcore > 10 heuristics
CSUQ > Computer System Usability Questionnaire > Lewis, 1995 > IBM > 19 questions
ASQ > After Scenario Questionnaire > Lewis, 1995 > IBM > 3 questions
PHUE > Practical Heuristics for Usability Evaluation > Perlman, 1997 > OSU > 13 heuristics
PUTQ > Purdue Usability Testing Questionnaire > Lin et al, 1997 > Purdue > 100 questions
USE > USE Questionnaire > Lund, 2001 > Sapient > 30 questions
Questionnaires have long been used to evaluate user interfaces (Root & Draper, 1983). Questionnaires have also long been used in electronic form (Perlman, 1985). For a handful of questionnaires specifically designed to assess aspects of usability, the validity and/or reliability have been established, including some in the following table, some of which are discussed in Chapter 6 of Tullis & Albert, 2008.
There are other questionnaires, including instruments from the HFRG (Human Factors Research Group):
SUMI: Software Usability Measurement Inventory
MUMMS: Measurement of Usability of Multi Media Software
WAMMI: Website Analysis and Measurement Inventory
FAQ: Frequently Asked Questions about Questionnaires in Usability Engineering (compiled by Jurek Kirakowski)
NetRaker (netraker.com) once provided online usability evaluation, with a free trial version. Now there are many online survey services, most offering free simple surveys or trial periods:
SurveyMonkey
Zoomerang Survey
QuestionPro
Free Online Surveys
WebSurveyor
AdvancedSurvey
KeySurvey
Note: Although these instruments are online, many are copyrighted and some require a fee for non-educational use. See the references at the end of each form for more information.
Source: http://hcibib.org/perlman/question.html
User Interface Usability Evaluation with Web-Based Questionnaires
Last updated: 2009-07-15 by Gary Perlman (hcibib.org/perlman/ perlman@acm.org)
Acronym > Instrument > Reference > Institution > Example
QUIS > Questionnaire for User Interface Satisfaction > Chin et al, 1988 > Maryland >27 questions
PUEU > Perceived Usefulness and Ease of Use > Davis, 1989 > IBM > 12 questions
NAU > Nielsen's Attributes of Usability > Nielsen, 1993 > Bellcore > 5 attributes
NHE > Nielsen's Heuristic Evaluation > Nielsen, 1993 > Bellcore > 10 heuristics
CSUQ > Computer System Usability Questionnaire > Lewis, 1995 > IBM > 19 questions
ASQ > After Scenario Questionnaire > Lewis, 1995 > IBM > 3 questions
PHUE > Practical Heuristics for Usability Evaluation > Perlman, 1997 > OSU > 13 heuristics
PUTQ > Purdue Usability Testing Questionnaire > Lin et al, 1997 > Purdue > 100 questions
USE > USE Questionnaire > Lund, 2001 > Sapient > 30 questions
Questionnaires have long been used to evaluate user interfaces (Root & Draper, 1983). Questionnaires have also long been used in electronic form (Perlman, 1985). For a handful of questionnaires specifically designed to assess aspects of usability, the validity and/or reliability have been established, including some in the following table, some of which are discussed in Chapter 6 of Tullis & Albert, 2008.
There are other questionnaires, including instruments from the HFRG (Human Factors Research Group):
SUMI: Software Usability Measurement Inventory
MUMMS: Measurement of Usability of Multi Media Software
WAMMI: Website Analysis and Measurement Inventory
FAQ: Frequently Asked Questions about Questionnaires in Usability Engineering (compiled by Jurek Kirakowski)
NetRaker (netraker.com) once provided online usability evaluation, with a free trial version. Now there are many online survey services, most offering free simple surveys or trial periods:
SurveyMonkey
Zoomerang Survey
QuestionPro
Free Online Surveys
WebSurveyor
AdvancedSurvey
KeySurvey
Note: Although these instruments are online, many are copyrighted and some require a fee for non-educational use. See the references at the end of each form for more information.
Source: http://hcibib.org/perlman/question.html
Tuesday, October 6, 2009
Oct 6 - SUMI - Software Usability Measurement Inventory
I found a sample SUMI form...........year 1993, 2000. There are 50 questions. Answers are: Agree, Undecided, Disagree.
http://sumi.ucc.ie/uksample.pdf
A good resource for SUMI.
http://sumi.ucc.ie/
http://www2.hf.faa.gov/workbenchtools/default.aspx?rPage=Tooldetails&subCatId=13&toolID=245
FAA website has info on the various Usability Questionnaires.
About SUMI
SUMI was developed on the project by the Human Factors Research Group (HFRG), University College, Cork.
This generic usability tool comprises a validated 50-item paper-based questionnaire in which respondents score each item on a three-point scale (i.e., agree, undecided, disagree).
SUMI measures software quality from the end user's point of view.
The questionnaire is designed to measure scales of:
1) Affect - the respondents emotional feelings towards the software (e.g., warm, happy).
2) Efficiency - the sense of the degree to which the software enables the task to be completed in a timely, effective and economical fashion.
3) Learnability - the feeling that it is relatively straightforward to become familiar with the software.
4) Helpfulness - the perception that the software communicates in a helpful way to assist in the resolution of difficulties.
5) Control - the feeling that the software responds to user inputs in a consistent way and that its workings can easily be internalized. (Source: Porteous, Kirakowski and Corbett, 1993).
Advantages of SUMI
1) SUMI provides an objective way of assessing user satisfaction.
2) Because SUMI scores are based on a standardized questionnaire, SUMI results can be compared across different systems.
3) SUMI is mentioned in the ISO 9241 standard as a recognized method of testing user satisfaction.
Disadvantages of SUMI
1) The results SUMI produces are only valid if the sample used is representative of the user population, if the questionnaire has been administered in the same way to all users sampled, and if the results are carefully interpreted.
2) Experience interpreting the results of SUMI outputs is essential
3) Questionnaires can only provide information of a general nature, they do not identify specific problems which can be related to designers.
http://sumi.ucc.ie/uksample.pdf
A good resource for SUMI.
http://sumi.ucc.ie/
http://www2.hf.faa.gov/workbenchtools/default.aspx?rPage=Tooldetails&subCatId=13&toolID=245
FAA website has info on the various Usability Questionnaires.
About SUMI
SUMI was developed on the project by the Human Factors Research Group (HFRG), University College, Cork.
This generic usability tool comprises a validated 50-item paper-based questionnaire in which respondents score each item on a three-point scale (i.e., agree, undecided, disagree).
SUMI measures software quality from the end user's point of view.
The questionnaire is designed to measure scales of:
1) Affect - the respondents emotional feelings towards the software (e.g., warm, happy).
2) Efficiency - the sense of the degree to which the software enables the task to be completed in a timely, effective and economical fashion.
3) Learnability - the feeling that it is relatively straightforward to become familiar with the software.
4) Helpfulness - the perception that the software communicates in a helpful way to assist in the resolution of difficulties.
5) Control - the feeling that the software responds to user inputs in a consistent way and that its workings can easily be internalized. (Source: Porteous, Kirakowski and Corbett, 1993).
Advantages of SUMI
1) SUMI provides an objective way of assessing user satisfaction.
2) Because SUMI scores are based on a standardized questionnaire, SUMI results can be compared across different systems.
3) SUMI is mentioned in the ISO 9241 standard as a recognized method of testing user satisfaction.
Disadvantages of SUMI
1) The results SUMI produces are only valid if the sample used is representative of the user population, if the questionnaire has been administered in the same way to all users sampled, and if the results are carefully interpreted.
2) Experience interpreting the results of SUMI outputs is essential
3) Questionnaires can only provide information of a general nature, they do not identify specific problems which can be related to designers.
Friday, September 25, 2009
Plan for week Sep 28 - Oct 3: Usability Questionnaire
For next week, Sep 28 till Oct 3, I plan to read on Usability Questionnaire.
Usability Questionnaire:
SUMI
QUIS
PSSUQ
SUS
ASQ
CUSI
Usability Questionnaire is also a method used for evaluating usability. In my readings, I have come accross this methodology. I know the Usability Questionnaire would contain a set of questions.
I would like to know more about this. Usability Questionnaire is a type of usability evaluation tool. My research title is "Usability Evaluation Tool for Mobile Learning Applications."
Hence, reading up (literature review) on the popular types of Usability Questionnaire would definitely be useful for my knowledge.
Usability Questionnaire:
SUMI
QUIS
PSSUQ
SUS
ASQ
CUSI
Usability Questionnaire is also a method used for evaluating usability. In my readings, I have come accross this methodology. I know the Usability Questionnaire would contain a set of questions.
I would like to know more about this. Usability Questionnaire is a type of usability evaluation tool. My research title is "Usability Evaluation Tool for Mobile Learning Applications."
Hence, reading up (literature review) on the popular types of Usability Questionnaire would definitely be useful for my knowledge.
Labels:
ASQ,
CUSI,
PSSUQ,
QUIS,
research planning,
SUMI,
SUS,
usability questionnaire
Friday, September 4, 2009
Sep 4 - Sauro & Kindlund, A Method to Standardize Usability Metrics Into a Single Score.


A Method to Standardize Usability Metrics Into a Single Score.
Jeff Sauro. PeopleSoft, Inc., Denver, Colorado USA. Jeff_Sauro@peoplesoft.com
Erika Kindlund. Intuit, Inc., Mountain View, California USA. Erika_Kindlund@intuit.com
CHI 2005, April 2–7, 2005, Portland, Oregon, USA
ABSTRACT
Current methods to represent system or task usability in a single metric do not include all the ANSI and ISO defined usability aspects: effectiveness, efficiency & satisfaction. We propose a method to simplify all the ANSI and ISO aspects of usability into a single, standardized and summated usability metric (SUM). In four data sets, totaling 1860 task observations, we show that these aspects of usability are correlated and equally weighted and present a quantitative model for usability. Using standardization techniques from Six Sigma, we propose a scalable process for standardizing disparate usability metrics and show how Principal Components Analysis can be used to establish appropriate weighting for a summated model.
SUM provides one continuous variable for summative usability evaluations that can be used in regression analysis, hypothesis testing and usability reporting.
In a summative usability evaluation, several metrics are available to the analyst for benchmarking the usability of a product. There is general agreement from the standards boards ANSI 2001[2] and ISO 9241 pt.11[18] as to what the dimensions of usability are (effectiveness, efficiency & satisfaction) and to a lesser extent which metrics are most commonly used to quantify those dimensions.
Effectiveness includes measures for completion rates and errors, efficiency is measured from time on task and satisfaction is summarized using any of a number of standardized satisfaction questionnaires (either collected on a task-by-task basis or at the end of a test session) [2],[18].
There have been attempts to derive a single measure for the construct of usability.
Babiker et al [3] derived a single metric for usability in hypertext systems using objective performance measures only.
Questionnaires such as the SUMI [22,23], PSSUQ[27], QUIS[7] and SUS[5] have users provide a subjective assessment of recently completed tasks or specific product issues and claim to derive a reliable and low-cost standardized measure of the overall usability or quality of use of a system.
While the authors of these questionnaires do not necessarily intend for the questionnaires to act as a single measure of usability (e.g. “QUIS was designed to assess users' subjective satisfaction with specific aspects of the human-computer interface” [7]), they are often used by practitioners as a way to measure usability with one number. Such usage is often not discouraged by the questionnaires’ instructions (e.g. “SUMI is the only commercially available questionnaire for the assessment of the usability of software” [22] and “The SUS scale is a Likert scale and yields a single number representing a composite measure of the overall usability of the system [5]”).
McGee uses a geometric averaging procedure (UME) to standardize ratios of participants’ subjective assessment ratings on tasks to derive a single score for task usability. His research identifies the potential for a standardized measure of usability to support usability comparisons across products, the same product over time, at lower levels of detail, and of tasks common to multiple products.
Lewis used a rank-based system when assessing competing products [25]. This approach creates a rank score comprised of both users’ objective performance measures and subjective assessment, but the resulting metric only represents a relative comparison between like-products with similar tasks.
My Comments: I may consider to use quantitative usability concept and a single, combined usability score.
METHOD
Four summative usability tests were conducted to collect the common metrics as described above (task completion, error counts, task times and satisfaction scores) as well as several other metrics as suggested in Dumas and Redish [11], and Nielsen [39].
For measuring satisfaction we created a questionnaire containing semantic distance scales with five points, similar to the ASQ created by Lewis [26] (see Table 5 below). The questionnaire included questions on task experience, ease of task, time on task, and overall task satisfaction.
The questionnaires were administered immediately after each task to improve accuracy [16]. The four usability tests were conducted in a controlled usability lab setting over a two-year period. Participants were asked to complete the tasks to the best of their ability and the administrator only intervened when the participant indicated they were done or gave up.
At the end of the test session, “post-test” satisfaction questions similar to those in SUMI and SUS that asked about overall product usability were given to users.
Data was collected from 129 total participants completing a total of 57 tasks. Participants varied in their application experience, gender, and industries.
RESULTS
Examining the Relationships between the Metrics
To attempt to combine the metrics into a single usability score we examined the relationship among the four primary variables for each task observation. We generated a correlation matrix with all four variables from all four data sets plus a combined data set containing data from all tests.
As can be seen in the lower right cell of Table 1, the Pearson Product Moment correlation coefficients between satisfaction and task completion are consistent with prior correlation analyses (that is, displaying moderate and significant correlations between .3 and .5) [26, 29]. What’s more, the positive correlation between subjective measures (satisfaction) and objective measures (time, errors and completion) are also consistent with Nielsen’s 1994 meta-analysis [38] (although the subjective measures were preferences instead of satisfaction in that study).
Frøkjær et al [12] earlier has made the case for including all aspects (effectiveness, efficiency and satisfaction) when measuring the usability of a system since it was found that these aspects did not always correlate in the data they reviewed. We agree with Frøkjær et al’s conclusion to measure all aspects of usability, however, not because they do not correlate with each other (our data clearly shows the opposite), but because each measure adds additional information not contained in the other measures.
Principal Component Analysis (PCA) was used as statitstical tool for analysis.
The goal then becomes standardizing the four variables (time, satisfaction, completion and errors).
STANDARDIZING USABILITY METRICS
To standardize each of the usability metrics we created a z-score type value or z-equivalent. For the continuous data (time and average satisfaction), we subtracted the mean value from a specification limit and divided by the standard deviation. For discrete data (completion rates and errors) we divided the unacceptable conditions (defects) by all opportunities for defects.
This method of standardization was adapted from the process sigma metric used in Six Sigma [4],[17], [43]. See Sauro & Kindlund [44] for a more detailed discussion on how to standardize these metrics from raw usability data.
Standardizing Task Completion
We can assume that all users want to successfully complete tasks, so a defect in task completion can be identified as an instance of a user failing a task. An opportunity for a defect in task completion is simply each instance of a user attempting a task. Therefore, we standardized task completion as the ratio of failed tasks to attempted tasks. This proportion of defects per opportunities has a corresponding z-equivalent that can be looked up in a standard normal table.
For example, a task completion rate of 80% would have the z-equivalent of .841.
Standardizing Error Rates
Each error instance is unique, yet all are associated with the more general “opportunity” to make an error in this component of the task. Once the task’s error opportunities have been identified, the z-equivalent can be calculated by dividing the total number of errors by the error opportunities. This proportion can be approximated using the standard normal deviate.
Standardizing Satisfaction Scores
As described in the Methods section, we used a post task questionnaire containing 5-point semantic distance scales with the end points labeled (e.g. 5:Very Easy to 1:Very Difficult). For the analysis we created a composite satisfaction score by averaging the responses from questions of overall ease, satisfaction and perceived task time (See Table 5) .
To standardize the composite score we looked to the literature for a logical specification limit. Prior research across numerous usability studies suggests that systems with “good-usability” typically have a mean rating of 4 on a 1-5 scale and 5.6 on a 1-7 scale [38]. Therefore we set the specification limit to 4. To arrive at a standardized z-equivalent for composite satisfaction we subtracted the average rating of a user’s satisfaction score from 4 and divided by the standard deviation.
Standardizing Task Times
Identifying ideal task times presents an interesting challenge: how long is too long for any given task?
Once the ideal task time has been set for each task, standardizing the task time involves subtracting the raw task time from the specification limit and dividing by the standard deviation to arrive at the z-equivalent.
Creating a Single, Standardized and Summated Usability Metric: SUM
We created a single, standardized and summated usability metric for each task by averaging together the four standardized values based on the equal weighting of the coefficients from the Principal Components Analysis.
CONCLUSION
A single, standardized and summated usability metric (SUM) cannot and should not take the place of diagnostic qualitative usability improvements typically found in formative evaluations. When a summative evaluation is used to quantitatively assess the “before and after” impact of design changes, the advantage of one score is in its ability to summarize the majority of variance in four integral summative usability measures.
SUM has two additional advantages. First it provides one continuous variable that can be used in regression analysis, hypothesis testing and in the same ways existing metrics are used to report usability. Second, a single metric based on logical specification limits provides an idea of how usable a task or product is without having to reference historical data. This score can then be used to report against other key business metrics.
References that I may want to read further in future:
1. Abran, A., Surya, W., Khelifi, A., Rilling, J., Seffah, A., Robert, F. (2003). Consolidating the ISO Usability Models. Paper presented at 11th annual International Software Quality Management Conference.
2. ANSI (2001). Common industry format for usability test reports (ANSI-NCITS 354-2001). Washington, DC: American National Standards Institute.
5. Brooke, J. (1996). SUS: A “quick and dirty” usability scale. In P. Jordan, B. Thomas, and B. Weerdmeester (Eds.), Usability Evaluation in Industry (pp.189-194). London: Taylor and Francis. See also http://www.cee.hw.ac.uk/~ph/sus.html
8. Cordes, R. E (1984). Application of Magnitude Estimation for Evaluating Software Ease of Use. In Gavriel Salvendy (Ed.) First USA-Japan Conference on Human Computer Interaction, Amsterdam: Elsevier Science Publishers.
12. Frøkjær, E., Hertzum, M., and Hornbæk, K. (2000) Measuring usability: are effectiveness, efficiency, and satisfaction really correlated? In Proc. CHI 2000, (pp.345-352). Washington, D.C.: ACM Press.
13. Gliem, J. and Gliem, R. (2003). Calculating, Interpreting, and Reporting Cronbach’s Alpha Reliability Coefficient for Likert-Type Scales. In 2003 Midwest Research to Practice Conference in Adult, Continuing and Community Education. Columbus, OH.
22. Kirakowski, J. (1996). The Software Usability Measurement Inventory: Background and usage. In P. Jordan, B. Thomas, and B. Weerdmeester (Eds.), Usability Evaluation in Industry (pp. 169-178). London, UK: Taylor and Francis. (Also, see http://www.ucc.ie/hfrg/questionnaires/sumi/index.html )
23. Kirakowski, J., and Corbett, M. (1993). SUMI: The Software Usability Measurement Inventory. British Journal of Educational Technology, 24, 210-212.
25. Lewis, J (1991) A Rank-Based Method for the Usability Comparison of Competing Products. In Proceedings of the Human Factors and Ergonomics Society 35th Annual Meeting San Francisco California (pp1312-1316).
26. Lewis, J. R. (1991). Psychometric evaluation of an after-scenario questionnaire for computer usability studies: The ASQ. SIGCHI Bulletin, 23, 78-81.
27. Lewis, J. R. (1992). Psychometric evaluation of the Post-Study System Usability Questionnaire: The PSSUQ. In Proceedings of the Human Factors Society 36th Annual Meeting (pp. 1259-1263). Atlanta, GA: Human Factors Society.
28. Lewis, J. R. (1993). IBM computer usability satisfaction questionnaires: Psychometric evaluation and instructions for use (Tech. Report 54.786). Boca Raton, FL: IBM Corp. http://drjim.0catch.com/usabqtr.pdf
29. Lewis, J. R. (1995). IBM computer usability satisfaction questionnaires: Psychometric evaluation and instructions for use. International Journal of Human-Computer Interaction, 7, 57-78.
32. McGee, M. (2003). Usability magnitude estimation. Proc. HFES, 47th Annual Meeting, (691--695).
33. McGee, M (2004). Master usability scaling: magnitude estimation and master scaling applied to usability measurement. In Proc. CHI 2004, (pp 335 - 342). Washington, D.C.: ACM Press.
38. Nielsen, J. and Levy, J. (1994) Measuring Usability: Preference vs. Performance. Communications of the ACM, 37, p. 66-76
42. Sauro, J. (2004) How long should a task take? Identifying Spec Limits for Task Times in Usability Tests. Retrieved September 13, 2004, from Measuring Usability Web site : http://measuringusability.com/time_specs.htm
43. Sauro, J. (2004) How Do You Calculate a Z-Score? Retrieved September 13, 2004, from Measuring Usability Web site: http://measuringusability.com/z_calc.htm
44. Sauro, J & Kindlund E. (In Press) Making Sense of Usability Metrics: Usability and Six Sigma, in Proceedings of the 14th Annual Conference of the Usability Professionals Association, Montreal, Canada
Tuesday, September 1, 2009
Sep 1 - Jokela et al, Methods for quantitative usability requirements: a case study on the development of the user interface of a mobile phone

Methods for quantitative usability requirements: a case study on the development of the user interface of a mobile phone.
Timo Jokela Æ Jussi Koivumaa Æ Jani Pirkola
Petri Salminen Æ Niina Kantola
Received: 3 February 2005 / Accepted: 4 May 2005 / Published online: 8 October 2005
Springer-Verlag London Limited 2005
Timo Jokela Æ Jussi Koivumaa Æ Jani Pirkola
Petri Salminen Æ Niina Kantola
Received: 3 February 2005 / Accepted: 4 May 2005 / Published online: 8 October 2005
Springer-Verlag London Limited 2005
Pers Ubiquit Comput (2006) 10: 345–355
DOI 10.1007/s00779-005-0050-7
T. Jokela (&) Æ N. Kantola. Oulu University, P.O. Box 3000, Oulu, Finland. E-mail: timo.jokela@oulu.fi E-mail: niina.kantola@oulu.fi
J. Koivumaa Æ J. Pirkola. Nokia, P.O. Box 50, 90571 Oulu, Finland. E-mail: jussi.koivumaa@nokia.com E-mail: jani.pirkola@nokia.com
P. Salminen. ValueFirst, Luuvantie 28, 02620 Espoo, Finland. E-mail: petri.salminen@valuefirst.fi
J. Koivumaa Æ J. Pirkola. Nokia, P.O. Box 50, 90571 Oulu, Finland. E-mail: jussi.koivumaa@nokia.com E-mail: jani.pirkola@nokia.com
P. Salminen. ValueFirst, Luuvantie 28, 02620 Espoo, Finland. E-mail: petri.salminen@valuefirst.fi
Abstract
Quantitative usability requirements are a critical but challenging, and hence an often neglected aspect of a usability engineering process. A case study is described where quantitative usability requirements played a key role in the development of a new user interface of a mobile phone. Within the practical constraints of the project, existing methods for determining usability requirements and evaluating the extent to which these are met, could not be applied as such, therefore tailored methods had to be developed. These methods and their applications are discussed.
Mobile phones have become a natural part of our everyday lives. Their user friendliness, termed usability, are increasingly in demand. Usability brings many benefits: users are able and willing to use the various features of the phone and the services supplied by the operators, the need for customer support decreases, and, above all, user satisfaction increases.
At the same time, designing is becoming increasingly challenging with the increasing number of functions and reduction of the size of the phones. Another challenge is the ever shortening life of the phones resulting in less time for development.
The practice of designing usable products is called usability engineering.1 The book User-centered system design by Donald Norman and Stephen Draper [1] is a pioneering work. John Gould and his colleagues also worked with usability methodologies in the 1980s [2]. Dennis Wixon and Karen Holtzblatt at Digital Equipment developed Contextual Inquiry and later on Contextual Design [3]; Carroll and Mack [4] were also early contributors. Later, various UCD methodologies were proposed e.g. by [5–10]. The standard ISO 13407 [11] is a widely used general reference for usability engineering.
The first activity is to identify users. Context of use analysis is about getting to know users: what the users’ goals are in relation to the product under development, what kind of tasks they do and in which contexts. User information is the basis for usability requirements where the target levels of the usability of the product under development are determined. A new product should lead to more efficient user tasks...
An essential part of the usability life-cycle is (quantitative) usability requirements, i.e. measurable usability targets for the interaction design [13–17]. As stated in [13]: ‘‘Without measurable usability specifications, there is no way to determine the usability needs of a product, or to measure whether or not the finished product fulfils those needs. If we cannot measure usability, we cannot have usability engineering’’.
In this article, our aim is to meet the research challenge posed by Wixon: we present the methods that we used in a real development context of a mobile phone UI, for the determination of quantitative usability requirements and the evaluation of the compliance with them.
Methods for quantitative usability requirements
There are two main activities related to quantitative usability requirements.
During the early phases of a development project, the usability requirements are determined (a in Fig. 1), and during the late phases, the usability of the product is evaluated against the requirements (b in Fig. 1).
Determining usability requirements can be further split into two activities: defining the usability attributes, and setting target values or the attributes.
In the evaluation, a measuring instrument is required.
Determining usability attributes
The main reference of usability is probably the definition of usability in ISO 9241-11: ‘‘the extent to which a product can be used by specified users to achieve specified goals with effectiveness, efficiency and satisfaction in a specified context of use’’ [19]. In brief, the definition means that usability requirements are based on measures of users performing tasks with the product to be developed.
– An example of an effectiveness measure is the percentage of users who can successfully complete a task.
– Efficiency can be measured by the mean time needed to successfully complete a task.
– User satisfaction can be measured with a questionnaire.
Usability requirements may include separate definitions of the target level (e.g. 90% of users can successfully complete a task) and the minimum acceptable level (e.g. 80% of users can successfully complete a task) [20].
Whiteside et al. [21] suggest that quantitative usability requirements be phrased at four levels: worst, planned, best and current.
Questionnaires measuring user satisfaction provide quantitative, though subjective usability metrics for related usability attributes.
Methods for determining usability targets
Possibly one of the most detailed guidelines for determining usability requirements is a six-step process by Wixon and Wilson [14].
In their process, relevant usability attributes are determined based on user profile and task analysis. Then the measuring instruments and measures are decided upon and a performance level is set for each attribute. T
hey agree with Whiteside et al. [21] that four performance levels can be set for each attribute
and determining the current level lays the foundation for setting other levels.
and determining the current level lays the foundation for setting other levels.
If the product is new, measurements for the current level can be attained, for example, from an existing manual system. Like Hix and Hartson [6], Wixon and Wilson [14] suggest that in the beginning two to three clear goals that focus on important and frequent tasks are enough and later, as the development teams accept the value of usability goals, more complex specifications can be generated.
Gould and Lewis [26] state that developing behavioural goals must cover at least three points.
Firstly, a description of the intended users must be given and the experimental participants should be agreed upon.
Secondly, the tasks to be performed and the circumstances in which they should be performed must be given.
The third point of the process is giving the measurement of interest, such as learning time and the criterion values to be achieved for each.
According to Nielsen [5] usability is associated with five attributes: learnability, efficiency, memorability, error and satisfaction. In usability goal setting, these attributes must be prioritised based on user and task analysis, and then operationalised and expressed in measurable ways.
Mayhew [7] introduces a nine-step procedure for setting usability goals. In her procedure qualitative usability goals are first identified and prioritized. Then those qualitative usability goals that are relatively high priority and seem easily quantifiable should be formulated to quantified goals.
For example the MUSiC methodology [28] aims to provide a comprehensive approach to the measurement of usability. It includes methods for specifying and measuring usability during design. One of the methods is the performance measurement method, which aims to provide a means of measuring two of the ISO 9241-11 standard usability components, i.e. effectiveness and efficiency.
My Comments: I think "quantitative usability" concept would be useful for my PhD research. Particularly relevant would be measurement of Effectiveness and Efficiency. Satisfaction would be more complex to be measured.
Methods for quantitative evaluation of usability
Whether the quantitative requirements have been met can be determind through a usability test.
Wixon et al. [14] define the term ‘‘test’’ as a broad term that encompasses any method for assessing whether goals have been achieved, like a formal laboratory test or a collection of satisfaction data through survey.
When evaluating usability, ISO 9241-11 [19] claims it is important that the context selected be representative. Evaluations can be done in the field in a real work situation or in laboratory settings in which the relevant aspects of the context of use are re-created in a representative and controlled way. A method that includes representative users performing typical, representative tasks is generally called usability testing.
Tasks that are done in usability testing provide an objective metric for the related usability attribute. Hix and Hartson [6] indicate that tasks must be very specifically worded in order to be the same for each participant. Tasks must also be specific, so that participants do not get sidetracked into irrelevant details during testing.
Wixon et al. [14] suggest that during the test, the tester should minimize the interaction with participants. Butler [33] describes his approach where ‘‘seven test users were given an introductory level problem to solve, then left alone with a copy of the user’s guide and a 3270 terminal logged onto the system.’’
User preference questionnaires provide a subjective metric for the related usability attribute such as ease of use or usefulness. Questionnaires are commonly built using Likert and semantic differential scales and are intended for use in various circumstances [32].
User preference questionnaires provide a subjective metric for the related usability attribute such as ease of use or usefulness. Questionnaires are commonly built using Likert and semantic differential scales and are intended for use in various circumstances [32].
There are a number of questionnaires available for quantitative usability evaluation, like SUMI [22], QUIS [24] and SUS [23]. Karat [34] states that questionnaires provide an easy and inexpensive method for obtaining measurement data on a system.
Usability can be quantitatively evaluated also with theory-based approaches such as GOMS and keystroke level model, KLM [35]. With GOMS, for example, total times can be predicted by associating times with each operator.
According to John [36], GOMS can also be used to predict how long it will take to learn a certain
task. With these quantitative predictions GOMS can be applied for example in a comparison between two systems.
The GOMS model also has its limitations. Preece et al. [37] suggest that GOMS can only really model computer-based tasks that involve a small set of highly routine data-entry type inputs. The model is not appropriate if errors occur.
KLM is a simplified version of GOMS.
task. With these quantitative predictions GOMS can be applied for example in a comparison between two systems.
The GOMS model also has its limitations. Preece et al. [37] suggest that GOMS can only really model computer-based tasks that involve a small set of highly routine data-entry type inputs. The model is not appropriate if errors occur.
KLM is a simplified version of GOMS.
My Comments: Read this article in detail...will enhance understanding of how they did the quantitative usability.
Discussion
Determining appropriate usability attributes and setting target values for them is a challenging task. Usability requirements should be defined so that they depict a ‘usable product’ as well as possible. .... It should be understood, however, that the appropriate set of attributes is heavily dependent on the product or application.
The determination of quantitative usability requirements and their evaluation should be distinguished. We propose that it is not necessary to know how to measure them exactly at the time of determining the requirements. An important role of usability requirements is that they give direction and vision to the user interface design.
This experience is shared by Wixon et al. [14]: ‘‘even if you do not test at all, designing with a clearly stated usability goal is preferable to designing toward a generic goal of ‘easy and intuitive’’’.
We encourage innovativeness in usability methods. It is seldom possible to use usability methods ideally. This article presents our innovations on the methods for determining and evaluating usability requirements. ..... The project context and the business case always have a major impact on the usability attributes.
Conclusion
We described a case study from a development project where the use of quantitative usability requirements was found useful.
We used methods that do not exactly follow the existing well-known usability methods. We believe that this is not a unique case: most industrial 5 development projects have specific constraints and limitations, and an ideal use of usability methods is not generally feasible. While we strongly recommend the use of measurable usability requirements, we do not propose our methods as a general solution. Clearly, each project has its specific features, and the usability methods should be selected and tailored based on the specific context of the project.
References that I may want to read further in future:
7. Mayhew DJ (1999) The usability engineering lifecycle. Morgan Kaufman, San Fancisco
14. Wixon D, Wilson C (1997) The usability engineering framework for product design and evaluation. In: Helander M, Landauer T, Prabhu P (eds) Handbook of human–computer
interaction. Elsevier, Amsterdam. pp 653–688
18. Wixon D (2003) Evaluating usability methods. Why the current literature fails the practitioner. Interactions 10(4):28–34
20. NIST (2004) Proposed industry format for usability requirements. Draft version 0.62
22. Kirakowski J, Corbett M (1993) SUMI: The software usability measurement inventory. Br J Educ Technol 24(3):210–212
23. Brooke J (1986) SUS — A ‘‘quick and dirty’’ usability scale. Digital Equipment Co. Ltd
24. Chin JP, Diehl VA, Norman KL (1988) Development of an instrument measuring user satisfaction of the human–computer interface. In: Proceedings of SIGCHI ‘88. New York
27. Dumas JS, Redish JC (1993)A practical guide to usability testing. Ablex Publishing Corporation, Norwood
28. Bevan N, Macleod M (1994) Usability measurement in context. Behav Inf Technol 13(1,2):132–145
29. Macleod M, Bowden R, Bevan N, Curson I (1997) The MUSiC performance measurement method. Behav Inf Technol 16(4,5):279–293
30. Maguire M (1998) RESPECT user-centred requirements handbook. Version 3.3. HUSAT Research Institute (now the Ergonomics and Saftety Research Institute, ESRI), Loughborough
University
14. Wixon D, Wilson C (1997) The usability engineering framework for product design and evaluation. In: Helander M, Landauer T, Prabhu P (eds) Handbook of human–computer
interaction. Elsevier, Amsterdam. pp 653–688
18. Wixon D (2003) Evaluating usability methods. Why the current literature fails the practitioner. Interactions 10(4):28–34
20. NIST (2004) Proposed industry format for usability requirements. Draft version 0.62
22. Kirakowski J, Corbett M (1993) SUMI: The software usability measurement inventory. Br J Educ Technol 24(3):210–212
23. Brooke J (1986) SUS — A ‘‘quick and dirty’’ usability scale. Digital Equipment Co. Ltd
24. Chin JP, Diehl VA, Norman KL (1988) Development of an instrument measuring user satisfaction of the human–computer interface. In: Proceedings of SIGCHI ‘88. New York
27. Dumas JS, Redish JC (1993)A practical guide to usability testing. Ablex Publishing Corporation, Norwood
28. Bevan N, Macleod M (1994) Usability measurement in context. Behav Inf Technol 13(1,2):132–145
29. Macleod M, Bowden R, Bevan N, Curson I (1997) The MUSiC performance measurement method. Behav Inf Technol 16(4,5):279–293
30. Maguire M (1998) RESPECT user-centred requirements handbook. Version 3.3. HUSAT Research Institute (now the Ergonomics and Saftety Research Institute, ESRI), Loughborough
University
31. Bevan N, Claridge N, Athousaki M, Maguire M, Catarci T, Matarazzo G, Raiss G (2002) Guide to specifying and evaluating usability as part of a contract, version1.0. PRUE project. Serco Usability Services, London, p 47
32. ANSI (2001) Common industry format for usability test reports. NCITS 354–2001
32. ANSI (2001) Common industry format for usability test reports. NCITS 354–2001
Labels:
GOMS,
Jokela,
Kantola,
KLM,
Koivumaa,
Pirkola,
quantitative usability,
QUIS,
Salminen,
SUMI,
SUS,
usability,
user interface,
wireless mobile device,
Wixon
Subscribe to:
Posts (Atom)