A usability study of a postconference online selfassessment program for health care professionals by Harmons, Eric M., M.A., Texas Woman's University, 2005 , 49 pages; AAT 1428775 My Interest: 1) Usability analysis/study. 2) QUIS – Questionnaire User Interface Survey. 3) Phone interview. 4) Post-conference online assessment. 5) Blackboard Academic Suite. Action: To read the Dissertation in the future. Motivation Continuing education (CE) conferences are widely used by professionals for the enhancement of clinical proficiency. Professionals view these types of educational activities as a vital means to enhance their professional practice and maintain clinical competency. The current environment of continuing education conferences revolves around a single-phase (didactic only) approach to learning. Current literature supports the notion of a multi-phased approach to enhance the learning experience for the lifelong learner. Research Goal The purpose of this study was a pilot usability analysis of a post-conference online selfassessment, developed within a Blackboard Academic Suite(TM) (Blackboard) classroom. Methodology The Questionnaire User Interaction Survey (QUIS) and phone interviews were utilized to assess the user interaction satisfaction with completing the self-assessment program. Results Discussion The overall response of participants to the system was a mean of 7.86 (scale from 1--9) across all the categories of usability. Based on the results, the Blackboard system is viewed by healthcare professionals as an effective medium for delivering an online self-assessment. Participants felt the self-assessment provided an effective tool to reflect on the presented material and direct their learning. Comments: Softcopy Dissertation is scanned version; cannot copy and paste; difficult to blog about it. |
Showing posts with label QUIS. Show all posts
Showing posts with label QUIS. Show all posts
Sunday, September 26, 2010
20100926 - Harmons, Usability Study of Postconference Online Selfassessment...
Friday, October 9, 2009
Oct 10 - Tullis & Stetson, A Comparison of Questionnaires for Assessing Website Usability
A Comparison of Questionnaires for Assessing Website Usability
Thomas S. Tullis, Fidelity Investments
Jacqueline N. Stetson, Fidelity Investments and Bentley College
Various questionnaires have been reported in the literature for assessing the perceived usability of an interactive system, e.g:
–Questionnaire for User Interface Satisfaction (QUIS) (1988)
–Computer System Usability Questionnaire (CSUQ) (1995)
–System Usability Scale (SUS) (1996)
A slightly different approach was taken by Microsoft with their "Product Reaction Cards"(2002)
And we have been using our own questionnaire for several years in our Usability Lab at Fidelity Investments
Problem
How well do these questionnaires apply to the assessment of Websites?
Do any of these questionnaires work well, as an adjunct to a usability test, with relatively small numbers of users?
Our Study
Limited ourselves to questionnaires in the published literature
–Did not include commercial services for evaluating website usability (e.g., WAMMI, RelevantView, NetRaker, Vividence).
We studied five questionnaires:
–SUS
–QUIS
–CSUQ
–Microsoft’s "Words"
–Our own questionnaire
Questionnaire #1: SUS
•Developed at Digital Equipment Corp.
•Consists of ten items.
•Adapted by replacing "system"with "website".
•Each item is a statement (positive or negative) and a rating on a five-point scale of "Strongly Disagree"to "Strongly Agree".
Questionnaire #2: QUIS
•Developed at the University of Maryland.
•Original questionnaire had 27 questions.
–We dropped 3 that did not seem relevant to Websites (e.g., "Remembering names and use of commands").
•"System"was replaced by "website"and term "screen"was replaced by "web page".
•Each question is a rating on a ten-point scale with appropriate anchors.
Questionnaire #3: CSUQ
•Developed at IBM.
•Composed of 19 questions.
•"System"or "computer system"was replaced by "website".
•Each question is a statement and a rating on a seven-point scale of "Strongly Disagree"to "Strongly Agree".
Questionnaire #4: Words
•Based on the 118 words used by Microsoft on their Product Reaction Cards.
–Some positive (e.g., "Convenient")
–Some negative (e.g., "Unattractive")
•Each word was presented with a check-box
–Users were asked to choose the words that best describe their interaction with the website.
–Could choose as many or as few words as they wished.
Questionnaire #5: Ours
•Developed ourselves and have been using for several years in our usability tests of websites.
•Composed of nine statements (e.g., "This website is visually appealing") to which the user responds on a seven-point scale from "Strongly Disagree"to "Strongly Agree".
•Points of the scale are numbered -3, -2, -1, 0, 1, 2, 3.
–Obvious neutral point at 0.
A Live Experiment!
•We’re going to compare two sites:
–CircuitCity.com
–Outpost.com
•Task 1: Your digital camera uses SmartMediacards. Find the least expensive external reader (USB) for your PC that will read them.
•Task 2: You do lots of hiking. Find the least
expensive personal GPS with map capability
and at least 8 MB of memory.
Method of Our Study
•Conducted entirely on our company Intranet.
•123 of our employees participated.
•Each participant was randomly assigned to one of the five questionnaire conditions.
•Each was asked to perform two tasks on each of two well-known personal financial information sites.
•Sites studied:
–Finance.Yahoo.com
–Kiplinger.com
–Hereafter referred to only as "Site 1"and
"Site 2". Don’t assume which is which.
•Tasks:
–Find the highest price in the past year for a share of.
–Find the mutual fund with the highest 3-year return.
•Order of presentation of the two sites was randomized.
•After completing (or at least attempting) the two tasks on a site, the user was presented with the questionnaire for their randomly selected condition.
•Each user completed the same questionnaire for both sites.
Data Analysis
•For each participant, an overall score was calculated for each website by averaging all of the ratings on the questionnaire that was used.
–All scales had been coded internally so that the "better"end corresponded to higher numbers.
–These were converted to percentages by dividing each score by the maximum score possible on that scale.
–For example, a rating of 3 on SUS was converted to a percentage by dividing that by 5 (the maximum score for SUS), giving a percentage of 60%.
•Special treatment for the "Words"condition since it did not involve rating scales:
–Before the study, we classified each of the words as being "Positive"or "Negative".
–Not grouped or identified as such to the participants.
–For each participant, an overall score was calculated by counting the total number of words that person selected and then dividing that number into the number of "Positive"words chosen.
–If someone selected 8 positive words and 10 words total, that yielded a score of 80%.
Results
•Calculated frequency distributions for the ratings, converted to percentages, for:
–Each questionnaire
–Both websites
Bar Charts are used to compare the results. See http://www.upassoc.org/usability_resources/conference/2004/UPA-2004-TullisStetson.pdf
My Comments: This study will be a good benchmark and reference for my research.
Results: Summary
•All five questionnaires showed that Site 1 was significantly preferred over Site 2 (p<.01). •The largest mean difference (74% vs. 38%) was found using the Words questionnaire, but this was also the questionnaire that yielded the greatest variability. Analysis of Sub-samples
•Next we analyzed randomly selected sub-samples of the data at size 6, 8, 10, 12, and 14.
–20 random samples for each size
•For each sample, t-test was conducted to determine whether the results showed that Site 1 was significantly better than Site 2 (the conclusion from the full dataset).
Analysis of Sub-samples
•Accuracy of the results increases as the sample size gets larger.
•With a sample size of only 6, all of the questionnaires yield accuracy of only 30-40%
–60-70% of the time, at that sample size, you would fail to find a significant difference between the two sites.
•Accuracy of some of the questionnaires increases quicker than others.
–SUS jumps up to about 75% accuracy at a size of 8.
Caveats
•Results were undoubtedly influenced by:
–The sites studied.
–The tasks used.
•We have only addressed the question of whether a given questionnaire was able to reliably distinguish between the ratings of one site vs. the other.
–Often you care more about how well the results help guide a redesign.
Conclusions
•One of the simplest questionnaires studied, SUS (with only 10 rating scales), yielded among the most reliable results across sample sizes.
–Also the only one whose questions all address different aspects of the user’s reaction to the website as a whole.
•For the conditions of this study, sample sizes of at least 12-14 participants are needed to get reasonably reliable results.
Source:
http://www.upassoc.org/usability_resources/conference/2004/UPA-2004-TullisStetson.pdf
Thomas S. Tullis, Fidelity Investments
Jacqueline N. Stetson, Fidelity Investments and Bentley College
Various questionnaires have been reported in the literature for assessing the perceived usability of an interactive system, e.g:
–Questionnaire for User Interface Satisfaction (QUIS) (1988)
–Computer System Usability Questionnaire (CSUQ) (1995)
–System Usability Scale (SUS) (1996)
A slightly different approach was taken by Microsoft with their "Product Reaction Cards"(2002)
And we have been using our own questionnaire for several years in our Usability Lab at Fidelity Investments
Problem
How well do these questionnaires apply to the assessment of Websites?
Do any of these questionnaires work well, as an adjunct to a usability test, with relatively small numbers of users?
Our Study
Limited ourselves to questionnaires in the published literature
–Did not include commercial services for evaluating website usability (e.g., WAMMI, RelevantView, NetRaker, Vividence).
We studied five questionnaires:
–SUS
–QUIS
–CSUQ
–Microsoft’s "Words"
–Our own questionnaire
Questionnaire #1: SUS
•Developed at Digital Equipment Corp.
•Consists of ten items.
•Adapted by replacing "system"with "website".
•Each item is a statement (positive or negative) and a rating on a five-point scale of "Strongly Disagree"to "Strongly Agree".
Questionnaire #2: QUIS
•Developed at the University of Maryland.
•Original questionnaire had 27 questions.
–We dropped 3 that did not seem relevant to Websites (e.g., "Remembering names and use of commands").
•"System"was replaced by "website"and term "screen"was replaced by "web page".
•Each question is a rating on a ten-point scale with appropriate anchors.
Questionnaire #3: CSUQ
•Developed at IBM.
•Composed of 19 questions.
•"System"or "computer system"was replaced by "website".
•Each question is a statement and a rating on a seven-point scale of "Strongly Disagree"to "Strongly Agree".
Questionnaire #4: Words
•Based on the 118 words used by Microsoft on their Product Reaction Cards.
–Some positive (e.g., "Convenient")
–Some negative (e.g., "Unattractive")
•Each word was presented with a check-box
–Users were asked to choose the words that best describe their interaction with the website.
–Could choose as many or as few words as they wished.
Questionnaire #5: Ours
•Developed ourselves and have been using for several years in our usability tests of websites.
•Composed of nine statements (e.g., "This website is visually appealing") to which the user responds on a seven-point scale from "Strongly Disagree"to "Strongly Agree".
•Points of the scale are numbered -3, -2, -1, 0, 1, 2, 3.
–Obvious neutral point at 0.
A Live Experiment!
•We’re going to compare two sites:
–CircuitCity.com
–Outpost.com
•Task 1: Your digital camera uses SmartMediacards. Find the least expensive external reader (USB) for your PC that will read them.
•Task 2: You do lots of hiking. Find the least
expensive personal GPS with map capability
and at least 8 MB of memory.
Method of Our Study
•Conducted entirely on our company Intranet.
•123 of our employees participated.
•Each participant was randomly assigned to one of the five questionnaire conditions.
•Each was asked to perform two tasks on each of two well-known personal financial information sites.
•Sites studied:
–Finance.Yahoo.com
–Kiplinger.com
–Hereafter referred to only as "Site 1"and
"Site 2". Don’t assume which is which.
•Tasks:
–Find the highest price in the past year for a share of
–Find the mutual fund with the highest 3-year return.
•Order of presentation of the two sites was randomized.
•After completing (or at least attempting) the two tasks on a site, the user was presented with the questionnaire for their randomly selected condition.
•Each user completed the same questionnaire for both sites.
Data Analysis
•For each participant, an overall score was calculated for each website by averaging all of the ratings on the questionnaire that was used.
–All scales had been coded internally so that the "better"end corresponded to higher numbers.
–These were converted to percentages by dividing each score by the maximum score possible on that scale.
–For example, a rating of 3 on SUS was converted to a percentage by dividing that by 5 (the maximum score for SUS), giving a percentage of 60%.
•Special treatment for the "Words"condition since it did not involve rating scales:
–Before the study, we classified each of the words as being "Positive"or "Negative".
–Not grouped or identified as such to the participants.
–For each participant, an overall score was calculated by counting the total number of words that person selected and then dividing that number into the number of "Positive"words chosen.
–If someone selected 8 positive words and 10 words total, that yielded a score of 80%.
Results
•Calculated frequency distributions for the ratings, converted to percentages, for:
–Each questionnaire
–Both websites
Bar Charts are used to compare the results. See http://www.upassoc.org/usability_resources/conference/2004/UPA-2004-TullisStetson.pdf
My Comments: This study will be a good benchmark and reference for my research.
Results: Summary
•All five questionnaires showed that Site 1 was significantly preferred over Site 2 (p<.01). •The largest mean difference (74% vs. 38%) was found using the Words questionnaire, but this was also the questionnaire that yielded the greatest variability. Analysis of Sub-samples
•Next we analyzed randomly selected sub-samples of the data at size 6, 8, 10, 12, and 14.
–20 random samples for each size
•For each sample, t-test was conducted to determine whether the results showed that Site 1 was significantly better than Site 2 (the conclusion from the full dataset).
Analysis of Sub-samples
•Accuracy of the results increases as the sample size gets larger.
•With a sample size of only 6, all of the questionnaires yield accuracy of only 30-40%
–60-70% of the time, at that sample size, you would fail to find a significant difference between the two sites.
•Accuracy of some of the questionnaires increases quicker than others.
–SUS jumps up to about 75% accuracy at a size of 8.
Caveats
•Results were undoubtedly influenced by:
–The sites studied.
–The tasks used.
•We have only addressed the question of whether a given questionnaire was able to reliably distinguish between the ratings of one site vs. the other.
–Often you care more about how well the results help guide a redesign.
Conclusions
•One of the simplest questionnaires studied, SUS (with only 10 rating scales), yielded among the most reliable results across sample sizes.
–Also the only one whose questions all address different aspects of the user’s reaction to the website as a whole.
•For the conditions of this study, sample sizes of at least 12-14 participants are needed to get reasonably reliable results.
Source:
http://www.upassoc.org/usability_resources/conference/2004/UPA-2004-TullisStetson.pdf
Oct 10 - HCIL, QUIS: The Questionnaire for User Interaction Satisfaction
QUIS: The Questionnaire for User Interaction Satisfaction
Subjective evaluation is an important component in the evaluation of workstation usability.
We have developed and standardized a general user evaluation instrument for interactive computer systems. The methods of psychological test construction were applied in order to ensure proper construct and empirical validity of the items and to assess their reliability. A hierarchical approach was taken in which overall usability was divided into subcomponents which constituted independent psychometric scales. For example, subcomponents include character readability, usefulness of online help, and meaningfulness of error messages.
Evaluation on these scales is assessed by user ratings of specific system attributes such as character definition, contrast, font, and spacing for the scale of character readability.
The purpose of the questionnaire is to:
1. guide in the design or redesign of systems,
2. give managers a tool for assessing potential areas of system improvement,
3. provide researchers with a validated instrument for conducting comparative evaluations, and
4. serve as a test instrument in usability labs. Validation studies continue to be run. It was recently shown that mean ratings are virtually the same for paper versus computer versions of the QUIS, but the computer version elicits more and longer open-ended comments.
The QUIS is licensed through the Office of Technology Liaison. Short and long paper versions are available as well as online versions that run in Windows and Macintosh environments, and now in HTML. The QUIS is currently licensed to dozens of usability labs and research centers around the world.
Related Papers:
Slaughter, L., Norman, K.L., Shneiderman, B. (March 1995) Assessing users' subjective satisfaction with the Information System for Youth Services (ISYS),VA Tech Proc. of Third Annual Mid-Atlantic Human Factors Conference (Blacksburg, VA, March 26-28, 1995) 164-170.CS-TR-3463, CAR-TR-768
Chin, J. P., Diehl, V. A, Norman, K. (Sept. 1987) Development of an instrument measuring user satisfaction of the human-computer interface, Proc. ACM CHI '88 (Washington, DC) 213-218. CS-TR-1926, CAR-TR-328
Participants:
Kent Norman, Department of Psychology
Ben Shneiderman, Computer Science
Ben Harper, Department of Psychology
Source: http://www.cs.umd.edu/hcil/quis/
Subjective evaluation is an important component in the evaluation of workstation usability.
We have developed and standardized a general user evaluation instrument for interactive computer systems. The methods of psychological test construction were applied in order to ensure proper construct and empirical validity of the items and to assess their reliability. A hierarchical approach was taken in which overall usability was divided into subcomponents which constituted independent psychometric scales. For example, subcomponents include character readability, usefulness of online help, and meaningfulness of error messages.
Evaluation on these scales is assessed by user ratings of specific system attributes such as character definition, contrast, font, and spacing for the scale of character readability.
The purpose of the questionnaire is to:
1. guide in the design or redesign of systems,
2. give managers a tool for assessing potential areas of system improvement,
3. provide researchers with a validated instrument for conducting comparative evaluations, and
4. serve as a test instrument in usability labs. Validation studies continue to be run. It was recently shown that mean ratings are virtually the same for paper versus computer versions of the QUIS, but the computer version elicits more and longer open-ended comments.
The QUIS is licensed through the Office of Technology Liaison. Short and long paper versions are available as well as online versions that run in Windows and Macintosh environments, and now in HTML. The QUIS is currently licensed to dozens of usability labs and research centers around the world.
Related Papers:
Slaughter, L., Norman, K.L., Shneiderman, B. (March 1995) Assessing users' subjective satisfaction with the Information System for Youth Services (ISYS),VA Tech Proc. of Third Annual Mid-Atlantic Human Factors Conference (Blacksburg, VA, March 26-28, 1995) 164-170.CS-TR-3463, CAR-TR-768
Chin, J. P., Diehl, V. A, Norman, K. (Sept. 1987) Development of an instrument measuring user satisfaction of the human-computer interface, Proc. ACM CHI '88 (Washington, DC) 213-218. CS-TR-1926, CAR-TR-328
Participants:
Kent Norman, Department of Psychology
Ben Shneiderman, Computer Science
Ben Harper, Department of Psychology
Source: http://www.cs.umd.edu/hcil/quis/
Oct 10 - Usabilitynet, Usability Questionnaires
My comments: This is a good summary/intro of several Usability Questionnaires.
It has been observed that questionnaires are the most frequently used tools for usability evaluation. This page is a list of usability questionnaire resources, extending the information presented on the questionnaires page of Usabilitynet.
SUMI
This is a mature questionnaire whose standardisation base and manual have been regularly updated. It is good for desktop products, but has also been used to evaluate command-and-control applications. It is a commercial product which comes complete with scoring and report generation software. It is designed and sold by the Human Factors Research Group at University College Cork.
WAMMI
This is a new questionnaire, designed to evaluate the quality of use of web sites. It is backed up by an extensive standardisation database, and it is purchased on a per report basis. It is the result of a joint development project by Jurek Kirakowski and Nigel Claridge.
SUS
This is a mature questionnaire, developed by John Brooke in 1986 and not published until years later. It is very robust and has been extensively used and adapted. It is public domain and nobody has published any standardisation data about it. Of all the public domain questionnaires, this is the most strongly recommended.
QUIS
This is a questionnaire developed by Kent Norman that has been modified many times to keep it current since its first appearance. It is commercially available and is championed by Ben Shneiderman in his book Designing the User Interface. Reading, MA: Addison-Wesley Publishing Co., 1998. Despite lack of standardisation and validation data, it has many adherents.
USE (see http://www.mindspring.com/~alund/USE/IntroductionToUse.html)
This questionnaire is still in development by Arnie Lund (last updated 11/11/98). It attempts to create a three-factor model of usability that can be applied to many situations. However, no reliability or validation data are presented. Public domain use is encouraged.
CSUQ
This is a well-designed questionnaire developed by Jim Lewis and it is public domain. It has excellent psychometric reliability properties but no standardisation base.
IsoNorm (in German only)
This questionnaire is designed to test the usability quality of software following the ISO 9241 part 10 principles. It is created by a team led by Jochim Puemper. Strong reliabilities are claimed for the sub-scales, although it appears there may be a strong inter-correlation between them as well. Downloads and an on-line version are available from the above URL, as well as articles about it (all in German.)
IsoMetrics
This questionnaire is produced by Guenter Gediga and his team. It is another attempt to produce a way of measuring ISO 9241 part 10, with reference to specific software features that may give rise to low usability data. It is therefore good both for summative and formative assessments. The questionnaire is well researched and detailed statistical information is given. Downloads of English and German versions are available. There is no standardisation base for it but it is public domain.
Questionnaire resources
Source: http://www.usabilitynet.org/tools/r_questionnaire.htm
It has been observed that questionnaires are the most frequently used tools for usability evaluation. This page is a list of usability questionnaire resources, extending the information presented on the questionnaires page of Usabilitynet.
SUMI
This is a mature questionnaire whose standardisation base and manual have been regularly updated. It is good for desktop products, but has also been used to evaluate command-and-control applications. It is a commercial product which comes complete with scoring and report generation software. It is designed and sold by the Human Factors Research Group at University College Cork.
WAMMI
This is a new questionnaire, designed to evaluate the quality of use of web sites. It is backed up by an extensive standardisation database, and it is purchased on a per report basis. It is the result of a joint development project by Jurek Kirakowski and Nigel Claridge.
SUS
This is a mature questionnaire, developed by John Brooke in 1986 and not published until years later. It is very robust and has been extensively used and adapted. It is public domain and nobody has published any standardisation data about it. Of all the public domain questionnaires, this is the most strongly recommended.
QUIS
This is a questionnaire developed by Kent Norman that has been modified many times to keep it current since its first appearance. It is commercially available and is championed by Ben Shneiderman in his book Designing the User Interface. Reading, MA: Addison-Wesley Publishing Co., 1998. Despite lack of standardisation and validation data, it has many adherents.
USE (see http://www.mindspring.com/~alund/USE/IntroductionToUse.html)
This questionnaire is still in development by Arnie Lund (last updated 11/11/98). It attempts to create a three-factor model of usability that can be applied to many situations. However, no reliability or validation data are presented. Public domain use is encouraged.
CSUQ
This is a well-designed questionnaire developed by Jim Lewis and it is public domain. It has excellent psychometric reliability properties but no standardisation base.
IsoNorm (in German only)
This questionnaire is designed to test the usability quality of software following the ISO 9241 part 10 principles. It is created by a team led by Jochim Puemper. Strong reliabilities are claimed for the sub-scales, although it appears there may be a strong inter-correlation between them as well. Downloads and an on-line version are available from the above URL, as well as articles about it (all in German.)
IsoMetrics
This questionnaire is produced by Guenter Gediga and his team. It is another attempt to produce a way of measuring ISO 9241 part 10, with reference to specific software features that may give rise to low usability data. It is therefore good both for summative and formative assessments. The questionnaire is well researched and detailed statistical information is given. Downloads of English and German versions are available. There is no standardisation base for it but it is public domain.
Questionnaire resources
Source: http://www.usabilitynet.org/tools/r_questionnaire.htm
Thursday, October 8, 2009
Oct 9 - HCIRN, Usability Questionnaires
Usability questionnaires intend to assess users' perception of the usability of a product.
A number of usability questionnaires have been developed or are currently under development:
ASQ (After-Scenario Questionnaire)
CSUQ (Computer System Usability Questionnaire)
CUSI (Computer User Satisfaction Inventory)
ErgoNorm Questionnaire
IsoMetrics
ISONORM 9241/10
MUMMS (Measuring the Usability of Multi-Media Systems)
PSSUQ (Post-Study System Usability Questionnaire)
PUEU (Perceived Usefulness and Ease of Use)
PUTQ (Purdue Usability Testing Questionnaire)
QUIS (Questionnaire for User Interaction Satisfaction)
SUMI (Software Usability Measurement Inventory)
SUS (System Usability Scale)
USE (Usefulness, Satisfaction, and Ease of Use)
WAMMI (Website Analysis and MeasureMent Inventory)
Some of these questionnaires are free, others are commercial and require a license to use them.
Two popular commercial questionnaires are QUIS (Questionnaire for User Interaction Satisfaction) and SUMI (Software Usability Measurement Inventory). Both provided analysis software with output in various formats.
Two popular free questionnaires are PSSUQ (Post-Study System Usability Questionnaire) and SUS (System Usability Scale).
Two of the questionnaires in the above list are related to the seven dialog principles described in ISO 9241 Part 10. IsoMetrics is directly based on these dialog principles and intends to measure the usability of a system along these dimensions. SUMI has five factors, at least four of which seem to directly correspond to dialog principles.
Global vs. Analytic Evaluation
Usability questionnaires generally provide only global measures of system usability. They do not provide any information about which aspects of the system contribute to the positive or negative evaluation. (Many authors use the terms formative and summative evaluation to distinguish between these types evaluation. However, based on a detailed analysis of these terms we believe that the terms analytic and global evaluation are more appropriate. See Formative and Summative Evaluation for more detail.)
Providing only a global measure is a major limitation of usability questionnaires. A number of questionnaires, therefore, attempt to facilitate analytic evaluation as well. One common way which nearly all usability questionnaires support is to provide a comments sections in which respondents can enter feedback in free form.
Some questionnaires also support some form of item analysis. A number of case studies using QUIS (Questionnaire for User Interaction Satisfaction) compared average item scores against the overall average score. However, this approach is questionable.
SUMI (Software Usability Measurement Inventory) uses a method called Item Consensual Analysis (ICA) to obtain additional information about usability problems. Item Consensual Analysis compares the response pattern of each item with its response pattern in the standardization database.
Assessing a Single System vs. Comparing Two or More Systems
Questionnaires can be used to assess a single system or to compare two or more systems. When using a questionnaire to assess a single system, the resulting score by itself is not very meaningful. What does it mean that your system has a score of 5.0 on a range from 1.0 to 7.0? It is better or worse than other comparable systems? To interpret a single score, you need a baseline to compare it with.
However, for a baseline to be useful, it has to be updated regularly. If the baseline score was 5.0 two years ago, it could be quite different now. Improvements to applications could have increased the baseline score to 6.0. Conversely, changes to applications may have increased their complexity which decreased the baseline score to 4.0.
Also, baselines need to be specific for the type of application. For example, new products, such as web browsers, may initially have a poorer score than established products, such as word processors. So a score of 5.0 might be poor compared to word processors and excellent compared to web browsers.
Finally, different baselines are needed for different countries. There might be cultural differences in the interpretation of specific questions as well as in the response tendency of the respondents. This makes it impossible to compare a score to a baseline obtained for another language version of the same questionnaire.
To our knowledge, to date only SUMI (Software Usability Measurement Inventory) provides such a standardization database.
Baselines are not required when comparing two or more systems. All usability questionnaires, provided they have sufficient reliability, are suitable for this. They can be used to measure changes in usability between versions of the same system or to assess differences between competing products.
Validity of Usability Questionnaires
Usability questionnaires can be interpreted in two ways. Firstly, they are a direct measure of users' perception of the usability of a system. Secondly, they are an indirect measure of the actual usability of a system.
...............
...............
One way to validate a questionnaire is to compare it to other questionnaires that claim to measure the same construct. However, few such studies have been conducted for usability questionnaires. Kirakowski (1994) reports two studies that directly compared CUSI (Computer User Satisfaction Inventory), SUS (System Usability Scale) and QUIS (Questionnaire for User Interaction Satisfaction) 5.0. Lucey (1991) found good correlations between the affect subscale of CUSI with SUS and overall QUIS. The correlations between the competence subscale of CUSI with SUS and QUIS were low. Wong & Rengger (1990) found a similar pattern.
Choosing a Usability Questionnaire
Which usability questionnaire should you choose? We have not had the opportunity to examine SUMI (Software Usability Measurement Inventory).
SUMI has never been published and is only available by purchasing the questionnaire package. However, it looks to be your best choice. It is a well-designed and extensively tested questionnaire. It is available in various languages and is the only usability questionnaire that has been standardized. Besides producing scores for a global scale and five subscales, detailed information about a system can be obtained by analyzing the response pattern of each item. The only disadvantage of SUMI is that it costs money. It is worth it, but HCI practitioners might find it difficult to convince management that the expense is justified.
Out of the freely available questionnaires, SUS (System Usability Scale) is probably your best choice. It does not have the standardization database of SUMI and it provides only an overall usability score. However, it is short and less susceptible to response bias than some other freely available questionnaires. SUS is an excellent instrument for comparing systems.
Compared to other evaluation methods in HCI, there is little literature available on usability questionnaires. There are also a number of weak spots in the published research on usability questionnaires. In particular, extensive validation studies that compare usability questionnaires with each other and with other usability evaluation methods are missing. The reason for this is not clear. Perhaps this is an indication that usability questionnaires are currently of little relevance in HCI practice.
Kirakowski, 1994 contains a good introduction to usability questionnaires.
Links
Questionnaires in Usability Engineering: A List of Frequently Asked Questions
By Jurek Kirakowski
http://www.ucc.ie/hfrg/resources/qfaq1.html
UsabilityNet > Questionnaire Resources
http://www.usabilitynet.org/tools/r_questionnaire.htm
A short overview of usability questionnaires.
Web-Based User Interface Evaluation with Questionnaires
By Gary Perlman
http://www.acm.org/~perlman/question.html
Includes online versions of a number of usability questionnaires.
Source: http://www.hcirn.com/atoz/atozu/usaques.php
A number of usability questionnaires have been developed or are currently under development:
ASQ (After-Scenario Questionnaire)
CSUQ (Computer System Usability Questionnaire)
CUSI (Computer User Satisfaction Inventory)
ErgoNorm Questionnaire
IsoMetrics
ISONORM 9241/10
MUMMS (Measuring the Usability of Multi-Media Systems)
PSSUQ (Post-Study System Usability Questionnaire)
PUEU (Perceived Usefulness and Ease of Use)
PUTQ (Purdue Usability Testing Questionnaire)
QUIS (Questionnaire for User Interaction Satisfaction)
SUMI (Software Usability Measurement Inventory)
SUS (System Usability Scale)
USE (Usefulness, Satisfaction, and Ease of Use)
WAMMI (Website Analysis and MeasureMent Inventory)
Some of these questionnaires are free, others are commercial and require a license to use them.
Two popular commercial questionnaires are QUIS (Questionnaire for User Interaction Satisfaction) and SUMI (Software Usability Measurement Inventory). Both provided analysis software with output in various formats.
Two popular free questionnaires are PSSUQ (Post-Study System Usability Questionnaire) and SUS (System Usability Scale).
Two of the questionnaires in the above list are related to the seven dialog principles described in ISO 9241 Part 10. IsoMetrics is directly based on these dialog principles and intends to measure the usability of a system along these dimensions. SUMI has five factors, at least four of which seem to directly correspond to dialog principles.
Global vs. Analytic Evaluation
Usability questionnaires generally provide only global measures of system usability. They do not provide any information about which aspects of the system contribute to the positive or negative evaluation. (Many authors use the terms formative and summative evaluation to distinguish between these types evaluation. However, based on a detailed analysis of these terms we believe that the terms analytic and global evaluation are more appropriate. See Formative and Summative Evaluation for more detail.)
Providing only a global measure is a major limitation of usability questionnaires. A number of questionnaires, therefore, attempt to facilitate analytic evaluation as well. One common way which nearly all usability questionnaires support is to provide a comments sections in which respondents can enter feedback in free form.
Some questionnaires also support some form of item analysis. A number of case studies using QUIS (Questionnaire for User Interaction Satisfaction) compared average item scores against the overall average score. However, this approach is questionable.
SUMI (Software Usability Measurement Inventory) uses a method called Item Consensual Analysis (ICA) to obtain additional information about usability problems. Item Consensual Analysis compares the response pattern of each item with its response pattern in the standardization database.
Assessing a Single System vs. Comparing Two or More Systems
Questionnaires can be used to assess a single system or to compare two or more systems. When using a questionnaire to assess a single system, the resulting score by itself is not very meaningful. What does it mean that your system has a score of 5.0 on a range from 1.0 to 7.0? It is better or worse than other comparable systems? To interpret a single score, you need a baseline to compare it with.
However, for a baseline to be useful, it has to be updated regularly. If the baseline score was 5.0 two years ago, it could be quite different now. Improvements to applications could have increased the baseline score to 6.0. Conversely, changes to applications may have increased their complexity which decreased the baseline score to 4.0.
Also, baselines need to be specific for the type of application. For example, new products, such as web browsers, may initially have a poorer score than established products, such as word processors. So a score of 5.0 might be poor compared to word processors and excellent compared to web browsers.
Finally, different baselines are needed for different countries. There might be cultural differences in the interpretation of specific questions as well as in the response tendency of the respondents. This makes it impossible to compare a score to a baseline obtained for another language version of the same questionnaire.
To our knowledge, to date only SUMI (Software Usability Measurement Inventory) provides such a standardization database.
Baselines are not required when comparing two or more systems. All usability questionnaires, provided they have sufficient reliability, are suitable for this. They can be used to measure changes in usability between versions of the same system or to assess differences between competing products.
Validity of Usability Questionnaires
Usability questionnaires can be interpreted in two ways. Firstly, they are a direct measure of users' perception of the usability of a system. Secondly, they are an indirect measure of the actual usability of a system.
...............
...............
One way to validate a questionnaire is to compare it to other questionnaires that claim to measure the same construct. However, few such studies have been conducted for usability questionnaires. Kirakowski (1994) reports two studies that directly compared CUSI (Computer User Satisfaction Inventory), SUS (System Usability Scale) and QUIS (Questionnaire for User Interaction Satisfaction) 5.0. Lucey (1991) found good correlations between the affect subscale of CUSI with SUS and overall QUIS. The correlations between the competence subscale of CUSI with SUS and QUIS were low. Wong & Rengger (1990) found a similar pattern.
Choosing a Usability Questionnaire
Which usability questionnaire should you choose? We have not had the opportunity to examine SUMI (Software Usability Measurement Inventory).
SUMI has never been published and is only available by purchasing the questionnaire package. However, it looks to be your best choice. It is a well-designed and extensively tested questionnaire. It is available in various languages and is the only usability questionnaire that has been standardized. Besides producing scores for a global scale and five subscales, detailed information about a system can be obtained by analyzing the response pattern of each item. The only disadvantage of SUMI is that it costs money. It is worth it, but HCI practitioners might find it difficult to convince management that the expense is justified.
Out of the freely available questionnaires, SUS (System Usability Scale) is probably your best choice. It does not have the standardization database of SUMI and it provides only an overall usability score. However, it is short and less susceptible to response bias than some other freely available questionnaires. SUS is an excellent instrument for comparing systems.
Compared to other evaluation methods in HCI, there is little literature available on usability questionnaires. There are also a number of weak spots in the published research on usability questionnaires. In particular, extensive validation studies that compare usability questionnaires with each other and with other usability evaluation methods are missing. The reason for this is not clear. Perhaps this is an indication that usability questionnaires are currently of little relevance in HCI practice.
Kirakowski, 1994 contains a good introduction to usability questionnaires.
Links
Questionnaires in Usability Engineering: A List of Frequently Asked Questions
By Jurek Kirakowski
http://www.ucc.ie/hfrg/resources/qfaq1.html
UsabilityNet > Questionnaire Resources
http://www.usabilitynet.org/tools/r_questionnaire.htm
A short overview of usability questionnaires.
Web-Based User Interface Evaluation with Questionnaires
By Gary Perlman
http://www.acm.org/~perlman/question.html
Includes online versions of a number of usability questionnaires.
Source: http://www.hcirn.com/atoz/atozu/usaques.php
Oct 8 - QUIS: QUestionnaire for Interaction Satisfaction
Questionnaire for User Interface Satisfaction
Please rate your satisfaction with the system.
Try to respond to all the items.
For items that are not applicable, use: NA
Likert Scale:
0
1
2
3
4
5
6
7
8
9
NA
OVERALL REACTION TO THE SOFTWARE
1. .......... > terrible > wonderful
2. .......... > difficult > easy
3. ........... > frustrating > satisfying
4. ........... > inadequate power > adequate power
5. ........... > dull > stimulating
6. ........... > rigid > flexible
SCREEN
7. Reading characters on the screen > hard > easy
8. Highlighting simplifies task > not at all > very much
9. Organization of information > confusing > very clear
10. Sequence of screens > confusing > very clear
TERMINOLOGY AND SYSTEM INFORMATION
11. Use of terms throughout system > inconsistent > consistent
12. Terminology related to task > never > always
13. Position of messages on screen > inconsistent > consistent
14. Prompts for input > confusing > clear
15. Computer informs about its progress > never > always
16. Error messages > unhelpful > helpful
LEARNING
17. Learning to operate the system > difficult > easy
18. Exploring new features by trial and error > difficult > easy
19. Remembering names and use of commands > difficult > easy
20. Performing tasks is straightforward > never > always
21. Help messages on the screen > unhelpful > helpful
22. Supplemental reference materials > confusing > clear
SYSTEM CAPABILITIES
23. System speed > too slow > fast enough
24. System reliability > unreliable > reliable
25. System tends to be > noisy > quiet
26. Correcting your mistakes > difficult > easy
27. Designed for all levels of users > never > always
List the most negative aspect(s):
1 ....................................................
2 ....................................................
3 ....................................................
List the most positive aspect(s):
1 ....................................................
2 ....................................................
3 ....................................................
Questionnaire for User Interface Satisfaction
Based on: Chin, J.P., Diehl, V.A., Norman, K.L. (1988) Development of an Instrument Measuring User Satisfaction of the Human-Computer Interface. ACM CHI'88 Proceedings, 213-218. ©1988 ACM. [Abstract] Copying without fee is permitted provided that the copies are not made or distributed for direct commercial advantage, and credit to the source is given. ©1986-1998 University of Maryland.
Source: http://hcibib.org/perlman/question.cgi?form=QUIS
Please rate your satisfaction with the system.
Try to respond to all the items.
For items that are not applicable, use: NA
Likert Scale:
0
1
2
3
4
5
6
7
8
9
NA
OVERALL REACTION TO THE SOFTWARE
1. .......... > terrible > wonderful
2. .......... > difficult > easy
3. ........... > frustrating > satisfying
4. ........... > inadequate power > adequate power
5. ........... > dull > stimulating
6. ........... > rigid > flexible
SCREEN
7. Reading characters on the screen > hard > easy
8. Highlighting simplifies task > not at all > very much
9. Organization of information > confusing > very clear
10. Sequence of screens > confusing > very clear
TERMINOLOGY AND SYSTEM INFORMATION
11. Use of terms throughout system > inconsistent > consistent
12. Terminology related to task > never > always
13. Position of messages on screen > inconsistent > consistent
14. Prompts for input > confusing > clear
15. Computer informs about its progress > never > always
16. Error messages > unhelpful > helpful
LEARNING
17. Learning to operate the system > difficult > easy
18. Exploring new features by trial and error > difficult > easy
19. Remembering names and use of commands > difficult > easy
20. Performing tasks is straightforward > never > always
21. Help messages on the screen > unhelpful > helpful
22. Supplemental reference materials > confusing > clear
SYSTEM CAPABILITIES
23. System speed > too slow > fast enough
24. System reliability > unreliable > reliable
25. System tends to be > noisy > quiet
26. Correcting your mistakes > difficult > easy
27. Designed for all levels of users > never > always
List the most negative aspect(s):
1 ....................................................
2 ....................................................
3 ....................................................
List the most positive aspect(s):
1 ....................................................
2 ....................................................
3 ....................................................
Questionnaire for User Interface Satisfaction
Based on: Chin, J.P., Diehl, V.A., Norman, K.L. (1988) Development of an Instrument Measuring User Satisfaction of the Human-Computer Interface. ACM CHI'88 Proceedings, 213-218. ©1988 ACM. [Abstract] Copying without fee is permitted provided that the copies are not made or distributed for direct commercial advantage, and credit to the source is given. ©1986-1998 University of Maryland.
Source: http://hcibib.org/perlman/question.cgi?form=QUIS
Oct 8 - Usability Questionnaires >>Perlman, User Interface Usability Evaluation with Web-Based Questionnaires
Usability Questionnaires
User Interface Usability Evaluation with Web-Based Questionnaires
Last updated: 2009-07-15 by Gary Perlman (hcibib.org/perlman/ perlman@acm.org)
Acronym > Instrument > Reference > Institution > Example
QUIS > Questionnaire for User Interface Satisfaction > Chin et al, 1988 > Maryland >27 questions
PUEU > Perceived Usefulness and Ease of Use > Davis, 1989 > IBM > 12 questions
NAU > Nielsen's Attributes of Usability > Nielsen, 1993 > Bellcore > 5 attributes
NHE > Nielsen's Heuristic Evaluation > Nielsen, 1993 > Bellcore > 10 heuristics
CSUQ > Computer System Usability Questionnaire > Lewis, 1995 > IBM > 19 questions
ASQ > After Scenario Questionnaire > Lewis, 1995 > IBM > 3 questions
PHUE > Practical Heuristics for Usability Evaluation > Perlman, 1997 > OSU > 13 heuristics
PUTQ > Purdue Usability Testing Questionnaire > Lin et al, 1997 > Purdue > 100 questions
USE > USE Questionnaire > Lund, 2001 > Sapient > 30 questions
Questionnaires have long been used to evaluate user interfaces (Root & Draper, 1983). Questionnaires have also long been used in electronic form (Perlman, 1985). For a handful of questionnaires specifically designed to assess aspects of usability, the validity and/or reliability have been established, including some in the following table, some of which are discussed in Chapter 6 of Tullis & Albert, 2008.
There are other questionnaires, including instruments from the HFRG (Human Factors Research Group):
SUMI: Software Usability Measurement Inventory
MUMMS: Measurement of Usability of Multi Media Software
WAMMI: Website Analysis and Measurement Inventory
FAQ: Frequently Asked Questions about Questionnaires in Usability Engineering (compiled by Jurek Kirakowski)
NetRaker (netraker.com) once provided online usability evaluation, with a free trial version. Now there are many online survey services, most offering free simple surveys or trial periods:
SurveyMonkey
Zoomerang Survey
QuestionPro
Free Online Surveys
WebSurveyor
AdvancedSurvey
KeySurvey
Note: Although these instruments are online, many are copyrighted and some require a fee for non-educational use. See the references at the end of each form for more information.
Source: http://hcibib.org/perlman/question.html
User Interface Usability Evaluation with Web-Based Questionnaires
Last updated: 2009-07-15 by Gary Perlman (hcibib.org/perlman/ perlman@acm.org)
Acronym > Instrument > Reference > Institution > Example
QUIS > Questionnaire for User Interface Satisfaction > Chin et al, 1988 > Maryland >27 questions
PUEU > Perceived Usefulness and Ease of Use > Davis, 1989 > IBM > 12 questions
NAU > Nielsen's Attributes of Usability > Nielsen, 1993 > Bellcore > 5 attributes
NHE > Nielsen's Heuristic Evaluation > Nielsen, 1993 > Bellcore > 10 heuristics
CSUQ > Computer System Usability Questionnaire > Lewis, 1995 > IBM > 19 questions
ASQ > After Scenario Questionnaire > Lewis, 1995 > IBM > 3 questions
PHUE > Practical Heuristics for Usability Evaluation > Perlman, 1997 > OSU > 13 heuristics
PUTQ > Purdue Usability Testing Questionnaire > Lin et al, 1997 > Purdue > 100 questions
USE > USE Questionnaire > Lund, 2001 > Sapient > 30 questions
Questionnaires have long been used to evaluate user interfaces (Root & Draper, 1983). Questionnaires have also long been used in electronic form (Perlman, 1985). For a handful of questionnaires specifically designed to assess aspects of usability, the validity and/or reliability have been established, including some in the following table, some of which are discussed in Chapter 6 of Tullis & Albert, 2008.
There are other questionnaires, including instruments from the HFRG (Human Factors Research Group):
SUMI: Software Usability Measurement Inventory
MUMMS: Measurement of Usability of Multi Media Software
WAMMI: Website Analysis and Measurement Inventory
FAQ: Frequently Asked Questions about Questionnaires in Usability Engineering (compiled by Jurek Kirakowski)
NetRaker (netraker.com) once provided online usability evaluation, with a free trial version. Now there are many online survey services, most offering free simple surveys or trial periods:
SurveyMonkey
Zoomerang Survey
QuestionPro
Free Online Surveys
WebSurveyor
AdvancedSurvey
KeySurvey
Note: Although these instruments are online, many are copyrighted and some require a fee for non-educational use. See the references at the end of each form for more information.
Source: http://hcibib.org/perlman/question.html
Tuesday, October 6, 2009
Oct 6 - QUIS - Questionnaire For Interaction Satisfaction
QUIS = QUestionnaire for Interaction Satisfaction
About QUIS
QUIS is a tool developed by the University of Maryland designed to assess users' subjective satisfaction with specific aspects of the human-computer interface.
The QUIS 7.0 is the current version.
It contains a demographic questionnaire, a measure of overall system satisfaction along six scales, and hierarchically organized measures of eleven specific interface factors (screen factors, terminology and system feedback, learning factors, system capabilities, technical manuals, on-line tutorials, multimedia, voice recognition, virtual environments, internet access, and software installation).
Each area measures the users' overall satisfaction with that facet of the interface, as well as the factors that make up that facet, on a 9-point scale.
The questionnaire is designed to be configured according to the needs of each interface analysis by including only the sections that are of interest to the user. (U. of Maryland)
QUIS can be purchased from the University of Maryland.
CONTACT:Office of Technology Commercialization
Attention: Doris Gillette 6200 Baltimore Ave., Suite 300, Riverdale, Maryland 20737, Office: (301) 403-2711, email: otl@umail.umd.edu
http://www2.hf.faa.gov/workbenchtools/default.aspxrPage=Tooldetails&subCatId=13&toolID=218
http://www.cs.umd.edu/hcil/ >> I got this reference. However, I opened up the link, I could not immediately see anything useful to me. Perhaps I may surf more slowly in the future.
More about QUIS
About the QUIS, version 7.0
The Questionnaire for User Interaction Satisfaction (QUIS) is a tool developed by a multi-disciplinary team of researchers in the Human-Computer Interaction Lab (HCIL) at the University of Maryland at College Park. The QUIS was designed to assess users' subjective satisfaction with specific aspects of the human-computer interface. The QUIS team successfully addressed the reliability and validity problems found in other satisfaction measures, creating a measure that is highly reliable across many types of interfaces.
The QUIS 7.0 is the current version. It contains a demographic questionnaire, a measure of overall system satisfaction along six scales, and hierarchically organized measures of nine specific interface factors (screen factors, terminology and system feedback, learning factors, system capabilities, technical manuals, on-line tutorials, multimedia, teleconferencing, and software installation). Each area measures the users' overall satisfaction with that facet of the interface, as well as the factors that make up that facet, on a 9-point scale. The questionnaire is designed to be configured according to the needs of each interface analysis by including only the sections that are of interest to the user.
In addition to English, the QUIS 7.0 is currently available in the following languages: German, Italian, Portuguese (Brazilian), and Spanish.
About QUIS
QUIS is a tool developed by the University of Maryland designed to assess users' subjective satisfaction with specific aspects of the human-computer interface.
The QUIS 7.0 is the current version.
It contains a demographic questionnaire, a measure of overall system satisfaction along six scales, and hierarchically organized measures of eleven specific interface factors (screen factors, terminology and system feedback, learning factors, system capabilities, technical manuals, on-line tutorials, multimedia, voice recognition, virtual environments, internet access, and software installation).
Each area measures the users' overall satisfaction with that facet of the interface, as well as the factors that make up that facet, on a 9-point scale.
The questionnaire is designed to be configured according to the needs of each interface analysis by including only the sections that are of interest to the user. (U. of Maryland)
QUIS can be purchased from the University of Maryland.
CONTACT:Office of Technology Commercialization
Attention: Doris Gillette 6200 Baltimore Ave., Suite 300, Riverdale, Maryland 20737, Office: (301) 403-2711, email: otl@umail.umd.edu
http://www2.hf.faa.gov/workbenchtools/default.aspxrPage=Tooldetails&subCatId=13&toolID=218
http://www.cs.umd.edu/hcil/ >> I got this reference. However, I opened up the link, I could not immediately see anything useful to me. Perhaps I may surf more slowly in the future.
More about QUIS
About the QUIS, version 7.0
The Questionnaire for User Interaction Satisfaction (QUIS) is a tool developed by a multi-disciplinary team of researchers in the Human-Computer Interaction Lab (HCIL) at the University of Maryland at College Park. The QUIS was designed to assess users' subjective satisfaction with specific aspects of the human-computer interface. The QUIS team successfully addressed the reliability and validity problems found in other satisfaction measures, creating a measure that is highly reliable across many types of interfaces.
The QUIS 7.0 is the current version. It contains a demographic questionnaire, a measure of overall system satisfaction along six scales, and hierarchically organized measures of nine specific interface factors (screen factors, terminology and system feedback, learning factors, system capabilities, technical manuals, on-line tutorials, multimedia, teleconferencing, and software installation). Each area measures the users' overall satisfaction with that facet of the interface, as well as the factors that make up that facet, on a 9-point scale. The questionnaire is designed to be configured according to the needs of each interface analysis by including only the sections that are of interest to the user.
In addition to English, the QUIS 7.0 is currently available in the following languages: German, Italian, Portuguese (Brazilian), and Spanish.
Friday, September 25, 2009
Plan for week Sep 28 - Oct 3: Usability Questionnaire
For next week, Sep 28 till Oct 3, I plan to read on Usability Questionnaire.
Usability Questionnaire:
SUMI
QUIS
PSSUQ
SUS
ASQ
CUSI
Usability Questionnaire is also a method used for evaluating usability. In my readings, I have come accross this methodology. I know the Usability Questionnaire would contain a set of questions.
I would like to know more about this. Usability Questionnaire is a type of usability evaluation tool. My research title is "Usability Evaluation Tool for Mobile Learning Applications."
Hence, reading up (literature review) on the popular types of Usability Questionnaire would definitely be useful for my knowledge.
Usability Questionnaire:
SUMI
QUIS
PSSUQ
SUS
ASQ
CUSI
Usability Questionnaire is also a method used for evaluating usability. In my readings, I have come accross this methodology. I know the Usability Questionnaire would contain a set of questions.
I would like to know more about this. Usability Questionnaire is a type of usability evaluation tool. My research title is "Usability Evaluation Tool for Mobile Learning Applications."
Hence, reading up (literature review) on the popular types of Usability Questionnaire would definitely be useful for my knowledge.
Labels:
ASQ,
CUSI,
PSSUQ,
QUIS,
research planning,
SUMI,
SUS,
usability questionnaire
Tuesday, September 1, 2009
Sep 1 - Jokela et al, Methods for quantitative usability requirements: a case study on the development of the user interface of a mobile phone

Methods for quantitative usability requirements: a case study on the development of the user interface of a mobile phone.
Timo Jokela Æ Jussi Koivumaa Æ Jani Pirkola
Petri Salminen Æ Niina Kantola
Received: 3 February 2005 / Accepted: 4 May 2005 / Published online: 8 October 2005
Springer-Verlag London Limited 2005
Timo Jokela Æ Jussi Koivumaa Æ Jani Pirkola
Petri Salminen Æ Niina Kantola
Received: 3 February 2005 / Accepted: 4 May 2005 / Published online: 8 October 2005
Springer-Verlag London Limited 2005
Pers Ubiquit Comput (2006) 10: 345–355
DOI 10.1007/s00779-005-0050-7
T. Jokela (&) Æ N. Kantola. Oulu University, P.O. Box 3000, Oulu, Finland. E-mail: timo.jokela@oulu.fi E-mail: niina.kantola@oulu.fi
J. Koivumaa Æ J. Pirkola. Nokia, P.O. Box 50, 90571 Oulu, Finland. E-mail: jussi.koivumaa@nokia.com E-mail: jani.pirkola@nokia.com
P. Salminen. ValueFirst, Luuvantie 28, 02620 Espoo, Finland. E-mail: petri.salminen@valuefirst.fi
J. Koivumaa Æ J. Pirkola. Nokia, P.O. Box 50, 90571 Oulu, Finland. E-mail: jussi.koivumaa@nokia.com E-mail: jani.pirkola@nokia.com
P. Salminen. ValueFirst, Luuvantie 28, 02620 Espoo, Finland. E-mail: petri.salminen@valuefirst.fi
Abstract
Quantitative usability requirements are a critical but challenging, and hence an often neglected aspect of a usability engineering process. A case study is described where quantitative usability requirements played a key role in the development of a new user interface of a mobile phone. Within the practical constraints of the project, existing methods for determining usability requirements and evaluating the extent to which these are met, could not be applied as such, therefore tailored methods had to be developed. These methods and their applications are discussed.
Mobile phones have become a natural part of our everyday lives. Their user friendliness, termed usability, are increasingly in demand. Usability brings many benefits: users are able and willing to use the various features of the phone and the services supplied by the operators, the need for customer support decreases, and, above all, user satisfaction increases.
At the same time, designing is becoming increasingly challenging with the increasing number of functions and reduction of the size of the phones. Another challenge is the ever shortening life of the phones resulting in less time for development.
The practice of designing usable products is called usability engineering.1 The book User-centered system design by Donald Norman and Stephen Draper [1] is a pioneering work. John Gould and his colleagues also worked with usability methodologies in the 1980s [2]. Dennis Wixon and Karen Holtzblatt at Digital Equipment developed Contextual Inquiry and later on Contextual Design [3]; Carroll and Mack [4] were also early contributors. Later, various UCD methodologies were proposed e.g. by [5–10]. The standard ISO 13407 [11] is a widely used general reference for usability engineering.
The first activity is to identify users. Context of use analysis is about getting to know users: what the users’ goals are in relation to the product under development, what kind of tasks they do and in which contexts. User information is the basis for usability requirements where the target levels of the usability of the product under development are determined. A new product should lead to more efficient user tasks...
An essential part of the usability life-cycle is (quantitative) usability requirements, i.e. measurable usability targets for the interaction design [13–17]. As stated in [13]: ‘‘Without measurable usability specifications, there is no way to determine the usability needs of a product, or to measure whether or not the finished product fulfils those needs. If we cannot measure usability, we cannot have usability engineering’’.
In this article, our aim is to meet the research challenge posed by Wixon: we present the methods that we used in a real development context of a mobile phone UI, for the determination of quantitative usability requirements and the evaluation of the compliance with them.
Methods for quantitative usability requirements
There are two main activities related to quantitative usability requirements.
During the early phases of a development project, the usability requirements are determined (a in Fig. 1), and during the late phases, the usability of the product is evaluated against the requirements (b in Fig. 1).
Determining usability requirements can be further split into two activities: defining the usability attributes, and setting target values or the attributes.
In the evaluation, a measuring instrument is required.
Determining usability attributes
The main reference of usability is probably the definition of usability in ISO 9241-11: ‘‘the extent to which a product can be used by specified users to achieve specified goals with effectiveness, efficiency and satisfaction in a specified context of use’’ [19]. In brief, the definition means that usability requirements are based on measures of users performing tasks with the product to be developed.
– An example of an effectiveness measure is the percentage of users who can successfully complete a task.
– Efficiency can be measured by the mean time needed to successfully complete a task.
– User satisfaction can be measured with a questionnaire.
Usability requirements may include separate definitions of the target level (e.g. 90% of users can successfully complete a task) and the minimum acceptable level (e.g. 80% of users can successfully complete a task) [20].
Whiteside et al. [21] suggest that quantitative usability requirements be phrased at four levels: worst, planned, best and current.
Questionnaires measuring user satisfaction provide quantitative, though subjective usability metrics for related usability attributes.
Methods for determining usability targets
Possibly one of the most detailed guidelines for determining usability requirements is a six-step process by Wixon and Wilson [14].
In their process, relevant usability attributes are determined based on user profile and task analysis. Then the measuring instruments and measures are decided upon and a performance level is set for each attribute. T
hey agree with Whiteside et al. [21] that four performance levels can be set for each attribute
and determining the current level lays the foundation for setting other levels.
and determining the current level lays the foundation for setting other levels.
If the product is new, measurements for the current level can be attained, for example, from an existing manual system. Like Hix and Hartson [6], Wixon and Wilson [14] suggest that in the beginning two to three clear goals that focus on important and frequent tasks are enough and later, as the development teams accept the value of usability goals, more complex specifications can be generated.
Gould and Lewis [26] state that developing behavioural goals must cover at least three points.
Firstly, a description of the intended users must be given and the experimental participants should be agreed upon.
Secondly, the tasks to be performed and the circumstances in which they should be performed must be given.
The third point of the process is giving the measurement of interest, such as learning time and the criterion values to be achieved for each.
According to Nielsen [5] usability is associated with five attributes: learnability, efficiency, memorability, error and satisfaction. In usability goal setting, these attributes must be prioritised based on user and task analysis, and then operationalised and expressed in measurable ways.
Mayhew [7] introduces a nine-step procedure for setting usability goals. In her procedure qualitative usability goals are first identified and prioritized. Then those qualitative usability goals that are relatively high priority and seem easily quantifiable should be formulated to quantified goals.
For example the MUSiC methodology [28] aims to provide a comprehensive approach to the measurement of usability. It includes methods for specifying and measuring usability during design. One of the methods is the performance measurement method, which aims to provide a means of measuring two of the ISO 9241-11 standard usability components, i.e. effectiveness and efficiency.
My Comments: I think "quantitative usability" concept would be useful for my PhD research. Particularly relevant would be measurement of Effectiveness and Efficiency. Satisfaction would be more complex to be measured.
Methods for quantitative evaluation of usability
Whether the quantitative requirements have been met can be determind through a usability test.
Wixon et al. [14] define the term ‘‘test’’ as a broad term that encompasses any method for assessing whether goals have been achieved, like a formal laboratory test or a collection of satisfaction data through survey.
When evaluating usability, ISO 9241-11 [19] claims it is important that the context selected be representative. Evaluations can be done in the field in a real work situation or in laboratory settings in which the relevant aspects of the context of use are re-created in a representative and controlled way. A method that includes representative users performing typical, representative tasks is generally called usability testing.
Tasks that are done in usability testing provide an objective metric for the related usability attribute. Hix and Hartson [6] indicate that tasks must be very specifically worded in order to be the same for each participant. Tasks must also be specific, so that participants do not get sidetracked into irrelevant details during testing.
Wixon et al. [14] suggest that during the test, the tester should minimize the interaction with participants. Butler [33] describes his approach where ‘‘seven test users were given an introductory level problem to solve, then left alone with a copy of the user’s guide and a 3270 terminal logged onto the system.’’
User preference questionnaires provide a subjective metric for the related usability attribute such as ease of use or usefulness. Questionnaires are commonly built using Likert and semantic differential scales and are intended for use in various circumstances [32].
User preference questionnaires provide a subjective metric for the related usability attribute such as ease of use or usefulness. Questionnaires are commonly built using Likert and semantic differential scales and are intended for use in various circumstances [32].
There are a number of questionnaires available for quantitative usability evaluation, like SUMI [22], QUIS [24] and SUS [23]. Karat [34] states that questionnaires provide an easy and inexpensive method for obtaining measurement data on a system.
Usability can be quantitatively evaluated also with theory-based approaches such as GOMS and keystroke level model, KLM [35]. With GOMS, for example, total times can be predicted by associating times with each operator.
According to John [36], GOMS can also be used to predict how long it will take to learn a certain
task. With these quantitative predictions GOMS can be applied for example in a comparison between two systems.
The GOMS model also has its limitations. Preece et al. [37] suggest that GOMS can only really model computer-based tasks that involve a small set of highly routine data-entry type inputs. The model is not appropriate if errors occur.
KLM is a simplified version of GOMS.
task. With these quantitative predictions GOMS can be applied for example in a comparison between two systems.
The GOMS model also has its limitations. Preece et al. [37] suggest that GOMS can only really model computer-based tasks that involve a small set of highly routine data-entry type inputs. The model is not appropriate if errors occur.
KLM is a simplified version of GOMS.
My Comments: Read this article in detail...will enhance understanding of how they did the quantitative usability.
Discussion
Determining appropriate usability attributes and setting target values for them is a challenging task. Usability requirements should be defined so that they depict a ‘usable product’ as well as possible. .... It should be understood, however, that the appropriate set of attributes is heavily dependent on the product or application.
The determination of quantitative usability requirements and their evaluation should be distinguished. We propose that it is not necessary to know how to measure them exactly at the time of determining the requirements. An important role of usability requirements is that they give direction and vision to the user interface design.
This experience is shared by Wixon et al. [14]: ‘‘even if you do not test at all, designing with a clearly stated usability goal is preferable to designing toward a generic goal of ‘easy and intuitive’’’.
We encourage innovativeness in usability methods. It is seldom possible to use usability methods ideally. This article presents our innovations on the methods for determining and evaluating usability requirements. ..... The project context and the business case always have a major impact on the usability attributes.
Conclusion
We described a case study from a development project where the use of quantitative usability requirements was found useful.
We used methods that do not exactly follow the existing well-known usability methods. We believe that this is not a unique case: most industrial 5 development projects have specific constraints and limitations, and an ideal use of usability methods is not generally feasible. While we strongly recommend the use of measurable usability requirements, we do not propose our methods as a general solution. Clearly, each project has its specific features, and the usability methods should be selected and tailored based on the specific context of the project.
References that I may want to read further in future:
7. Mayhew DJ (1999) The usability engineering lifecycle. Morgan Kaufman, San Fancisco
14. Wixon D, Wilson C (1997) The usability engineering framework for product design and evaluation. In: Helander M, Landauer T, Prabhu P (eds) Handbook of human–computer
interaction. Elsevier, Amsterdam. pp 653–688
18. Wixon D (2003) Evaluating usability methods. Why the current literature fails the practitioner. Interactions 10(4):28–34
20. NIST (2004) Proposed industry format for usability requirements. Draft version 0.62
22. Kirakowski J, Corbett M (1993) SUMI: The software usability measurement inventory. Br J Educ Technol 24(3):210–212
23. Brooke J (1986) SUS — A ‘‘quick and dirty’’ usability scale. Digital Equipment Co. Ltd
24. Chin JP, Diehl VA, Norman KL (1988) Development of an instrument measuring user satisfaction of the human–computer interface. In: Proceedings of SIGCHI ‘88. New York
27. Dumas JS, Redish JC (1993)A practical guide to usability testing. Ablex Publishing Corporation, Norwood
28. Bevan N, Macleod M (1994) Usability measurement in context. Behav Inf Technol 13(1,2):132–145
29. Macleod M, Bowden R, Bevan N, Curson I (1997) The MUSiC performance measurement method. Behav Inf Technol 16(4,5):279–293
30. Maguire M (1998) RESPECT user-centred requirements handbook. Version 3.3. HUSAT Research Institute (now the Ergonomics and Saftety Research Institute, ESRI), Loughborough
University
14. Wixon D, Wilson C (1997) The usability engineering framework for product design and evaluation. In: Helander M, Landauer T, Prabhu P (eds) Handbook of human–computer
interaction. Elsevier, Amsterdam. pp 653–688
18. Wixon D (2003) Evaluating usability methods. Why the current literature fails the practitioner. Interactions 10(4):28–34
20. NIST (2004) Proposed industry format for usability requirements. Draft version 0.62
22. Kirakowski J, Corbett M (1993) SUMI: The software usability measurement inventory. Br J Educ Technol 24(3):210–212
23. Brooke J (1986) SUS — A ‘‘quick and dirty’’ usability scale. Digital Equipment Co. Ltd
24. Chin JP, Diehl VA, Norman KL (1988) Development of an instrument measuring user satisfaction of the human–computer interface. In: Proceedings of SIGCHI ‘88. New York
27. Dumas JS, Redish JC (1993)A practical guide to usability testing. Ablex Publishing Corporation, Norwood
28. Bevan N, Macleod M (1994) Usability measurement in context. Behav Inf Technol 13(1,2):132–145
29. Macleod M, Bowden R, Bevan N, Curson I (1997) The MUSiC performance measurement method. Behav Inf Technol 16(4,5):279–293
30. Maguire M (1998) RESPECT user-centred requirements handbook. Version 3.3. HUSAT Research Institute (now the Ergonomics and Saftety Research Institute, ESRI), Loughborough
University
31. Bevan N, Claridge N, Athousaki M, Maguire M, Catarci T, Matarazzo G, Raiss G (2002) Guide to specifying and evaluating usability as part of a contract, version1.0. PRUE project. Serco Usability Services, London, p 47
32. ANSI (2001) Common industry format for usability test reports. NCITS 354–2001
32. ANSI (2001) Common industry format for usability test reports. NCITS 354–2001
Labels:
GOMS,
Jokela,
Kantola,
KLM,
Koivumaa,
Pirkola,
quantitative usability,
QUIS,
Salminen,
SUMI,
SUS,
usability,
user interface,
wireless mobile device,
Wixon
Subscribe to:
Posts (Atom)