Development of usability questionnaires for electronic mobile products and decision making methods by Ryu, Young Sam, Ph.D., Virginia Polytechnic Institute and State University, 2005 , 206 pages; AAT 3177322 My Interest: 1) MPUQ – Mobile Phone Usability Questionnaire. 2) How to select usability dimensions & questionnaire items. 3) Redundancy & Relevancy analysis. 4) How to test/validate MPUQ. Action: To read the Dissertation in future. Problem Statement As the growth of rapid prototyping techniques shortens the development life cycle of software and electronic products, usability inquiry methods can play a more significant role during the development life cycle, diagnosing usability problems and providing metrics for making comparative decisions. A need has been realized for questionnaires tailored to the evaluation of electronic mobile products, wherein usability is dependent on both hardware and software as well as the emotional appeal and aesthetic integrity of the design. Research Goal This research followed a systematic approach to develop a new questionnaire tailored to measure the usability of electronic mobile products. The Mobile Phone Usability Questionnaire (MPUQ) developed throughout this series of studies evaluates the usability of mobile phones for the purpose of making decisions among competing variations in the end-user market, alternatives of prototypes during the development process, and evolving versions during an iterative design process. In addition, the questionnaire can serve as a tool for identifying diagnostic information to improve specific usability dimensions and related interface elements. Methodology Employing the refined MPUQ, decision making models were developed using Analytic Hierarchy Process ( AHP) and linear regression analysis. Next, a new group of representative mobile users was Finally, a case study of comparative usability evaluations was performed to validate the MPUQ and models. A computerized support tool was developed to perform redundancy and relevancy analyses for the selection of appropriate questionnaire items. The weighted geometric mean was used to combine multiple numbers of matrices from pairwise comparison based on decision makers' consistency ratio values for AHP. The AHP and regression models provided important usability dimensions so that mobile device usability practitioners can simply focus on the interface elements related to the decisive usability dimensions in order to improve the usability of mobile products. The AHP model could predict the users' decision based on a descriptive model of purchasing the best product slightly but not significantly better than other evaluation methods. Mobile Phone Usability Questionnaire (MPUQ) Except for memorability, the MPUQ embraced the dimensions included in the other well-known usability definitions and almost all criteria covered by the existing usability questionnaires. In addition, MPUQ incorporated new criteria, such as pleasurability and specific tasks performance. 2.1. Subjective Usability Assessment 2.1.1. Definitions and Perspectives of Usability 2.1.2. Usability Measurements 2.1.3. Subjective Measurements of Usability 2.1.4. Usability Questionnaires 2.1.4.1. Definition of Questionnaire 2.1.4.2. Questionnaires and Usability Research 2.2. Mobile Device Usability 2.2.1. Definition of Electronic Mobile Products 2.2.2. Usability Concept for Mobile Device 2.2.3. Mobile Device Interfaces 2.2.4. Usability Testing for Mobile Device 3. PHASE I : DEVELOPMENT OF ITEMS AND CONTENT VALIDITY 3.1. Need for a New Scale 3.2. Study 1: Conceptualization and Development of Initial Items Pool 3.2.1. Conceptualization 3.2.2. Survey on Usability Dimensions and Criteria 3.2.2.1. Usability Dimensions by Early Studies 3.2.2.2. Usability Dimensions in Existing Usability Questionnaires 3.2.2.3. Usability Dimensions for Consumer Products 3.2.2.4. Items from a Usability Questionnaire for a Specific Product 3.2.3. Creation of an Items Pool 3.2.4. Choice of Format 3.3. Study 2: Subjective Usability Assessment Support Tool and Item Judgment 3.3.1. Method 3.3.1.1. Design 3.3.1.2. Equipment 3.3.1.3. Participants 3.3.1.4. Procedure 3.3.2. Result 3.3.2.1. Part 1. Redundancy Analysis 3.3.2.2. Part 2. Relevancy Analysis 3.3.3. Discussion 3.4. Outcome of Studies 1 and 2 4. PHASE II : REFINING QUESTIONNAIRE 4.1. Study 3: Questionnaire Item Analysis 4.1.1. Method 4.1.1.1. Design 4.1.1.2. Participants 4.1.1.3. Procedure 4.1.2. Results 4.1.2.1. User Information 4.1.2.2. Factor Analysis 4.1.2.3. Scale Reliability 4.1.2.4. Known-group Validity 4.1.3. Discussion 4.1.3.1. Eliminated Questionnaire Items 4.1.3.2. Normative Patterns 4.1.3.3. Limitations 4.2. Outcome of Study 3 6. PHASE IV : VALIDATION OF MODELS 6.1. Study 6: Comparative Evaluation with the Models 6.1.1. Method 6.1.1.1. Design 6.1.1.2. Equipment 6.1.1.3. Participants 6.1.1.4. Procedure 6.1.2. Results 6.1.2.1. Mean Rankings 6.1.2.2. Preference Data Format 6.1.2.3. Friedman Test for Minimalist 6.1.2.4. Friedman Test for Voice/Text Fanatics 6.1.2.5. Comparisons Among the Methods 6.1.2.6. Important Usability Dimensions 6.1.3. Discussion 6.1.3.1. Implication of Each Evaluation Method 6.1.3.2. PSSUQ and the MPUQ 6.1.3.3. Validity of MPUQ 6.1.3.4. Usability and Actual Purchase 6.1.3.5. Limitations 6.2. Outcome of Study 6 |
Showing posts with label PSSUQ. Show all posts
Showing posts with label PSSUQ. Show all posts
Friday, September 24, 2010
20100925 - Ryu, ..Usability Questionnaires for Electronic Mobile Products...
Friday, October 9, 2009
Oct 9 - PSSUQ: Post-Study System Usability Questionnaire - Lewis of IBM
The Post-Study System Usability Questionnaire (PSSUQ)
The Post-Study System Usability Questionnaire (PSSUQ) is currently a 19-item instrument for assessing user satisfaction with system usability. (See the appendix for a copy of the questionnaire items.)
Participants need more time to complete the PSSUQ than the ASQ (about 10 minutes to complete the PSSUQ), but only complete it once, at the end of a usability study. Completing the PSSUQ allows participants to provide an overall evaluation of the system they used.
After the 48 participants in the office-applications usability study (Lewis, Henry, & Mack, 1990) completed all the scenarios, they rated their system with the PSSUQ. This data allowed preliminary psychometric evaluation of the PSSUQ (Lewis, 1992b).
This earlier version of the PSSUQ (Lewis, 1992b) had only 18 items, with the items in a different order than shown in the appendix. Recently, a series of investigations using decision support systems revealed a common set of five system characteristics associated with usability by several different user groups (Doug Antonelli, personal communication, January 5, 1991). The original 18-item PSSUQ addressed four of these five system characteristics. The 19-item version of the PSSUQ contains an additional item to cover the fifth of these five system characteristics.
Item Construction
The items are 7-point graphic scales, anchored at the end points with the terms "Strongly agree" for 1, "Strongly disagree" for 7, and a "Not applicable" (N/A) point outside the scale.
Item Selection
A group of usability evaluators selected the items on the basis of their comprehensive content regarding hypothesized constituents of usability. For example, the items assess such system characteristics as ease of use, ease of learning, simplicity, effectiveness, information, and the user interface.
Psychometric Evaluation
Factor analysis.
The scree plot for an exploratory principal factors analysis of the PSSUQ data indicated that a 3-factor solution was appropriate (see Figure 2), so the overall scale defined by the full set of items contained three subscales. Table 5 shows the varimax-rotated factor pattern, revealing the structure of the subscales. Bold type in Table 5 highlights factor loadings that exceeded .5. Items that loaded highly on two factors were ambiguous regarding the appropriate subscale of which they should be a component, so they did not become a component of any subscale. (See the appendix to examine the content of these items.)
One of the most difficult tasks following this type of exploratory factor analysis is naming the factors. After considering a number of alternatives, a group of human factors engineers named the factors (and their corresponding subscales) System Usefulness (SYSUSE), Information Quality (INFOQUAL), and Interface Quality (INTERQUAL). These three factors account for 87% of the variability in the data.
Reliability.
Coefficient alpha analyses showed that the reliability of the overall summative scale (OVERALL) was .97, and ranged from .91 to .96 for the three subscales (SYSUSE=.96, INFOQUAL=.91, and INTERQUAL=.91). Therefore, the overall scale and the three subscales have excellent reliability.
Validity.
Correlation analyses support the validity of the scales. The OVERALL scale correlated highly with the sum of the ASQ ratings that participants gave after completing each scenario (r(20)=.80, p=.0001). OVERALL also correlated significantly with the percentage of successful scenario completion (r(29)=-.40, p=.026). The SYSUSE (r(36)=-.40, p=.006) and INTERQUAL (r(35)=-.29, p=.08) correlated with the percentage of successful scenario completion.
Sensitivity.
In the sensitivity ANOVAs, the overall scale and all three subscales indicated significant differences among the user groups (OVERALL: F(2,29)=4.35, p=.02; SYSUSE: F(2,36)=6.9, p=.003; INFOQUAL: F(2,33)=3.68, p=.04; INTERQUAL: F(2,33)=3.74, p=.03). INFOQUAL showed a significant system effect (F(2,33)=3.18, p=.05).
Discussion
These findings have limited generalizability because the sample size for the factor analysis was relatively small. The usual recommendation would be 90 participants for this questionnaire.
However, the factor analysis and reliability analyses suggest that it is reasonable to define three subscales from this set of items. The PSSUQ has reasonable concurrent validity when compared with successful scenario completion rates and the ASQ scores. The overall scale and the subscales are reasonably sensitive.
The evidence provided sufficient justification to use the PSSUQ to measure user satisfaction with system usability in usability studies, but also suggested that it would be prudent to collect more data in different circumstances to extend the generalizability of the findings.
My Comments: This would be helpful if I eventually select to develop Usability Questionnaire as usability evaluation tool.
IBM Computer Usability Satisfaction Questionnaires: Psychometric Evaluation and Instructions for Use
Technical Report 54.786
James R. Lewis
Human Factors Group
Boca Raton, FL
Source: http://drjim.0catch.com/usabqtr.pdf
The Post-Study System Usability Questionnaire (PSSUQ) is currently a 19-item instrument for assessing user satisfaction with system usability. (See the appendix for a copy of the questionnaire items.)
Participants need more time to complete the PSSUQ than the ASQ (about 10 minutes to complete the PSSUQ), but only complete it once, at the end of a usability study. Completing the PSSUQ allows participants to provide an overall evaluation of the system they used.
After the 48 participants in the office-applications usability study (Lewis, Henry, & Mack, 1990) completed all the scenarios, they rated their system with the PSSUQ. This data allowed preliminary psychometric evaluation of the PSSUQ (Lewis, 1992b).
This earlier version of the PSSUQ (Lewis, 1992b) had only 18 items, with the items in a different order than shown in the appendix. Recently, a series of investigations using decision support systems revealed a common set of five system characteristics associated with usability by several different user groups (Doug Antonelli, personal communication, January 5, 1991). The original 18-item PSSUQ addressed four of these five system characteristics. The 19-item version of the PSSUQ contains an additional item to cover the fifth of these five system characteristics.
Item Construction
The items are 7-point graphic scales, anchored at the end points with the terms "Strongly agree" for 1, "Strongly disagree" for 7, and a "Not applicable" (N/A) point outside the scale.
Item Selection
A group of usability evaluators selected the items on the basis of their comprehensive content regarding hypothesized constituents of usability. For example, the items assess such system characteristics as ease of use, ease of learning, simplicity, effectiveness, information, and the user interface.
Psychometric Evaluation
Factor analysis.
The scree plot for an exploratory principal factors analysis of the PSSUQ data indicated that a 3-factor solution was appropriate (see Figure 2), so the overall scale defined by the full set of items contained three subscales. Table 5 shows the varimax-rotated factor pattern, revealing the structure of the subscales. Bold type in Table 5 highlights factor loadings that exceeded .5. Items that loaded highly on two factors were ambiguous regarding the appropriate subscale of which they should be a component, so they did not become a component of any subscale. (See the appendix to examine the content of these items.)
One of the most difficult tasks following this type of exploratory factor analysis is naming the factors. After considering a number of alternatives, a group of human factors engineers named the factors (and their corresponding subscales) System Usefulness (SYSUSE), Information Quality (INFOQUAL), and Interface Quality (INTERQUAL). These three factors account for 87% of the variability in the data.
Reliability.
Coefficient alpha analyses showed that the reliability of the overall summative scale (OVERALL) was .97, and ranged from .91 to .96 for the three subscales (SYSUSE=.96, INFOQUAL=.91, and INTERQUAL=.91). Therefore, the overall scale and the three subscales have excellent reliability.
Validity.
Correlation analyses support the validity of the scales. The OVERALL scale correlated highly with the sum of the ASQ ratings that participants gave after completing each scenario (r(20)=.80, p=.0001). OVERALL also correlated significantly with the percentage of successful scenario completion (r(29)=-.40, p=.026). The SYSUSE (r(36)=-.40, p=.006) and INTERQUAL (r(35)=-.29, p=.08) correlated with the percentage of successful scenario completion.
Sensitivity.
In the sensitivity ANOVAs, the overall scale and all three subscales indicated significant differences among the user groups (OVERALL: F(2,29)=4.35, p=.02; SYSUSE: F(2,36)=6.9, p=.003; INFOQUAL: F(2,33)=3.68, p=.04; INTERQUAL: F(2,33)=3.74, p=.03). INFOQUAL showed a significant system effect (F(2,33)=3.18, p=.05).
Discussion
These findings have limited generalizability because the sample size for the factor analysis was relatively small. The usual recommendation would be 90 participants for this questionnaire.
However, the factor analysis and reliability analyses suggest that it is reasonable to define three subscales from this set of items. The PSSUQ has reasonable concurrent validity when compared with successful scenario completion rates and the ASQ scores. The overall scale and the subscales are reasonably sensitive.
The evidence provided sufficient justification to use the PSSUQ to measure user satisfaction with system usability in usability studies, but also suggested that it would be prudent to collect more data in different circumstances to extend the generalizability of the findings.
My Comments: This would be helpful if I eventually select to develop Usability Questionnaire as usability evaluation tool.
IBM Computer Usability Satisfaction Questionnaires: Psychometric Evaluation and Instructions for Use
Technical Report 54.786
James R. Lewis
Human Factors Group
Boca Raton, FL
Source: http://drjim.0catch.com/usabqtr.pdf
Thursday, October 8, 2009
Oct 9 - Lewis, IBM Computer Usability Satisfaction Questionnaires: Psychometric Evaluation and Instructions for Use
IBM Computer Usability Satisfaction Questionnaires: Psychometric Evaluation and Instructions for Use
ABSTRACT
This paper describes recent research in subjective usability measurement at IBM. The focus of the research was the application of psychometric methods to the development and evaluation of questionnaires that measure user satisfaction with system usability. The primary goals of this paper are to (1) discuss the psychometric characteristics of four IBM questionnaires that measure user satisfaction with computer system usability, and (2) provide the questionnaires, with administration and scoring instructions. Usability practitioners can use these questionnaires with confidence to help them measure users' satisfaction with the usability of computer systems.
Introduction
Customers want usable products, and developers strive to produce them. It follows that an important part of modern product engineering, both hardware and software, must be the measurement of usability.
Measuring usability is particularly difficult because usability is not a unidimensional product or user characteristic, but emerges as a multidimensional characteristic in the context of users performing tasks with a product in a specific environment (Bevan, Kirakowski, & Maissel, 1991; Shackel, 1984).
However, if you are unable to measure usability, how can you judge your product against your competitors', or even your own previous versions of the product?
Subjective and Objective Evaluation
Most usability evaluations gather both subjective and objective quantitative data in the context of realistic scenarios-of-use, as well as descriptions of the problems representative participants have trying to complete the scenarios.
Subjective data are measures of participants' opinions or attitudes concerning their perception of usability.
Objective data are measures of participants' performance (such as scenario completion time and successful scenario completion rate).
Objective usability measures include, but are not limited to, scenario completion time, successful scenario completion rate, and time spent recovering from errors (Whiteside, Bennett, & Holtzblatt, 1988). Subjective usability measures are usually responses to Likert-type questionnaire items that assess user attitude concerning attributes such as system ease-of-use and interface likeability (Alty, 1992).
Most usability evaluators collect both objective and subjective data.
Research Focus
The focus of this research was the application of psychometric methods to the development and evaluation of standard questionnaires to assess subjective usability.
The goal of psychometrics is to establish the quality of psychological measures (Nunnally, 1978). Is a measure reliable in the sense that it is consistent? Given a reliable measure, is it valid (measures the intended attribute)? Finally, is the measure appropriately sensitive to experimental manipulations?
Psychometrics is a well-developed field, but usability researchers have only recently used these methods to develop and evaluate questionnaires to assess usability (Sweeney & Dillon, 1987).
In contrast to other recent computer-user satisfaction questionnaires (Chin, Diehl, & Norman, 1988; Kirakowski & Dillon, 1988; LaLomia & Sidowski, 1990) the IBM questionnaires are specifically for use in the context of scenario-based usability testing (Lewis, 1991a; Lewis, 1991b; Lewis, 1991c; Lewis, 1992b; Lewis, Henry, & Mack, 1990), although additional research has indicated that one may be useful as an instrument for field evaluation (Lewis, 1992a). Usability practitioners can use these questionnaires to enhance their current usability methods. (The four IBM questionnaires appear in the appendix.)
Before describing the psychometric properties of the IBM questionnaires, I will briefly review the relevant elements of psychometric practice. (For a comprehensive discussion of psychometrics, see Nunnally, 1978.)
----------
ASQ
PSQ
PSSUQ
CSUQ
-----------
General Discussion
Although user satisfaction with system usability is only one component of the multifaceted construct of usability (Bevan et al., 1991), it is a very important component in many situations.
It is especially important when a primary design goal is user satisfaction.
This paper has described the psychometric qualities of four questionnaires that assess user satisfaction with system usability: the ASQ, PSQ, PSSUQ and CSUQ.
The ASQ and PSQ are both after-scenario questionnaires, intended for use in a scenario-based usability testing situation. They contain essentially the same items, but the ASQ uses a 7-point scale and the PSQ uses a 5-point scale.
Using data from very different scenario-based usability studies (one a study of software office applications, the other a study of printers), their factor analyses, validity analyses, and sensitivity analyses were virtually identical. Obtaining the same results in different settings with different user groups provides strong evidence that these results are generalizable, and the questionnaires have wide applicability. Because the ASQ has substantially better reliability than the PSQ, usability practitioners should use the ASQ rather than the PSQ as their after-scenario questionnaire.
The PSSUQ and CSUQ are both overall satisfaction questionnaires. The PSSUQ items are appropriate for a usability testing situation, and the CSUQ items are appropriate for a field testing situation. Otherwise, the questionnaires are identical.
The psychometric evaluations of the PSSUQ (using data from a usability study) and the CSUQ (using data from a mail survey) were virtually identical. As with the after-scenario questionnaires, this consistency provides strong evidence of generalizability of results and wide applicability of the questionnaires.
Because these questionnaires have acceptable psychometric properties, usability practitioners can use them with confidence as standardized measurements of satisfaction for usability studies and tests (ASQ, PSSUQ) or field research (CSUQ).
(Practitioners should note that nothing prevents the addition of items to these questionnaires if a particular situation suggests the need. However, using these questionnaires as the foundation for special-purpose questionnaires ensures that practitioners can score the scales and subscales from the questionnaires, maintaining the advantages of standardized measurement.)
Standardized satisfaction measurements offer many advantages to the usability practitioner (Nunnally, 1978). Specifically, standardized measurements provide:
1 Objectivity.
A standardized measurement supports objectivity because it allows usability practitioners to independently verify the measurement statements of other practitioners.
2 Quantification.
Standardized measurements allow practitioners to report results in finer detail than they could using only personal judgment. Standardization also permits practitioners to use powerful methods of mathematics and statistics to better understand their results (Nunnally, 1978).
3 Communication.
It is easier for practitioners to communicate effectively when standardized measures are available. Inadequate efficiency and fidelity of communication in any field is an impediment to progress.
4 Economy.
Developing standardized measures requires a substantial amount of work. However, once developed, they are economical. There is rarely any need to re-evaluate standardized measures.
5 Scientific generalization.
Scientific generalization is at the heart of scientific work. Standardization is essential for assessing the generalization of results.
Conclusion
In conclusion, these questionnaires should be valuable additions to the repertoire of techniques that usability practitioners apply in the design and evaluation of computer systems.
IBM Computer Usability Satisfaction Questionnaires: Psychometric Evaluation and Instructions for Use
Technical Report 54.786
James R. Lewis
Human Factors Group
Boca Raton, FL
Source: http://drjim.0catch.com/usabqtr.pdf
ABSTRACT
This paper describes recent research in subjective usability measurement at IBM. The focus of the research was the application of psychometric methods to the development and evaluation of questionnaires that measure user satisfaction with system usability. The primary goals of this paper are to (1) discuss the psychometric characteristics of four IBM questionnaires that measure user satisfaction with computer system usability, and (2) provide the questionnaires, with administration and scoring instructions. Usability practitioners can use these questionnaires with confidence to help them measure users' satisfaction with the usability of computer systems.
Introduction
Customers want usable products, and developers strive to produce them. It follows that an important part of modern product engineering, both hardware and software, must be the measurement of usability.
Measuring usability is particularly difficult because usability is not a unidimensional product or user characteristic, but emerges as a multidimensional characteristic in the context of users performing tasks with a product in a specific environment (Bevan, Kirakowski, & Maissel, 1991; Shackel, 1984).
However, if you are unable to measure usability, how can you judge your product against your competitors', or even your own previous versions of the product?
Subjective and Objective Evaluation
Most usability evaluations gather both subjective and objective quantitative data in the context of realistic scenarios-of-use, as well as descriptions of the problems representative participants have trying to complete the scenarios.
Subjective data are measures of participants' opinions or attitudes concerning their perception of usability.
Objective data are measures of participants' performance (such as scenario completion time and successful scenario completion rate).
Objective usability measures include, but are not limited to, scenario completion time, successful scenario completion rate, and time spent recovering from errors (Whiteside, Bennett, & Holtzblatt, 1988). Subjective usability measures are usually responses to Likert-type questionnaire items that assess user attitude concerning attributes such as system ease-of-use and interface likeability (Alty, 1992).
Most usability evaluators collect both objective and subjective data.
Research Focus
The focus of this research was the application of psychometric methods to the development and evaluation of standard questionnaires to assess subjective usability.
The goal of psychometrics is to establish the quality of psychological measures (Nunnally, 1978). Is a measure reliable in the sense that it is consistent? Given a reliable measure, is it valid (measures the intended attribute)? Finally, is the measure appropriately sensitive to experimental manipulations?
Psychometrics is a well-developed field, but usability researchers have only recently used these methods to develop and evaluate questionnaires to assess usability (Sweeney & Dillon, 1987).
In contrast to other recent computer-user satisfaction questionnaires (Chin, Diehl, & Norman, 1988; Kirakowski & Dillon, 1988; LaLomia & Sidowski, 1990) the IBM questionnaires are specifically for use in the context of scenario-based usability testing (Lewis, 1991a; Lewis, 1991b; Lewis, 1991c; Lewis, 1992b; Lewis, Henry, & Mack, 1990), although additional research has indicated that one may be useful as an instrument for field evaluation (Lewis, 1992a). Usability practitioners can use these questionnaires to enhance their current usability methods. (The four IBM questionnaires appear in the appendix.)
Before describing the psychometric properties of the IBM questionnaires, I will briefly review the relevant elements of psychometric practice. (For a comprehensive discussion of psychometrics, see Nunnally, 1978.)
----------
ASQ
PSQ
PSSUQ
CSUQ
-----------
General Discussion
Although user satisfaction with system usability is only one component of the multifaceted construct of usability (Bevan et al., 1991), it is a very important component in many situations.
It is especially important when a primary design goal is user satisfaction.
This paper has described the psychometric qualities of four questionnaires that assess user satisfaction with system usability: the ASQ, PSQ, PSSUQ and CSUQ.
The ASQ and PSQ are both after-scenario questionnaires, intended for use in a scenario-based usability testing situation. They contain essentially the same items, but the ASQ uses a 7-point scale and the PSQ uses a 5-point scale.
Using data from very different scenario-based usability studies (one a study of software office applications, the other a study of printers), their factor analyses, validity analyses, and sensitivity analyses were virtually identical. Obtaining the same results in different settings with different user groups provides strong evidence that these results are generalizable, and the questionnaires have wide applicability. Because the ASQ has substantially better reliability than the PSQ, usability practitioners should use the ASQ rather than the PSQ as their after-scenario questionnaire.
The PSSUQ and CSUQ are both overall satisfaction questionnaires. The PSSUQ items are appropriate for a usability testing situation, and the CSUQ items are appropriate for a field testing situation. Otherwise, the questionnaires are identical.
The psychometric evaluations of the PSSUQ (using data from a usability study) and the CSUQ (using data from a mail survey) were virtually identical. As with the after-scenario questionnaires, this consistency provides strong evidence of generalizability of results and wide applicability of the questionnaires.
Because these questionnaires have acceptable psychometric properties, usability practitioners can use them with confidence as standardized measurements of satisfaction for usability studies and tests (ASQ, PSSUQ) or field research (CSUQ).
(Practitioners should note that nothing prevents the addition of items to these questionnaires if a particular situation suggests the need. However, using these questionnaires as the foundation for special-purpose questionnaires ensures that practitioners can score the scales and subscales from the questionnaires, maintaining the advantages of standardized measurement.)
Standardized satisfaction measurements offer many advantages to the usability practitioner (Nunnally, 1978). Specifically, standardized measurements provide:
1 Objectivity.
A standardized measurement supports objectivity because it allows usability practitioners to independently verify the measurement statements of other practitioners.
2 Quantification.
Standardized measurements allow practitioners to report results in finer detail than they could using only personal judgment. Standardization also permits practitioners to use powerful methods of mathematics and statistics to better understand their results (Nunnally, 1978).
3 Communication.
It is easier for practitioners to communicate effectively when standardized measures are available. Inadequate efficiency and fidelity of communication in any field is an impediment to progress.
4 Economy.
Developing standardized measures requires a substantial amount of work. However, once developed, they are economical. There is rarely any need to re-evaluate standardized measures.
5 Scientific generalization.
Scientific generalization is at the heart of scientific work. Standardization is essential for assessing the generalization of results.
Conclusion
In conclusion, these questionnaires should be valuable additions to the repertoire of techniques that usability practitioners apply in the design and evaluation of computer systems.
IBM Computer Usability Satisfaction Questionnaires: Psychometric Evaluation and Instructions for Use
Technical Report 54.786
James R. Lewis
Human Factors Group
Boca Raton, FL
Source: http://drjim.0catch.com/usabqtr.pdf
Oct 9 - Post-Study System Usability Questionnaire (PSSUQ) - Lewis of IBM
The Post-Study System Usability Questionnaire (PSSUQ)
Administration and Scoring.
Give the PSSUQ to participants after they have completed all the scenarios in a usability study.
You can calculate four scores from the responses to the PSSUQ items:
* the overall satisfaction score (OVERALL),
* system usefulness (SYSUSE),
* information quality (INFOQUAL) and
* interface quality (INTERQUAL).
Because research on an alternative form of the PSSUQ (the Computer System Usability Questionnaire, or CSUQ) confirmed and clarified (and slightly modified) the factor structure of the questionnaire, refer to Appendix Table 1 in the next section of this appendix for the current scoring rules of the PSSUQ.
Instructions and Items.
The questionnaire's instructions and items are:
This questionnaire, which starts on the following page, gives you an opportunity to tell us your reactions to the system you used. Your responses will help us understand what aspects of the system you are particularly concerned about and the aspects that satisfy you.
To as great a degree as possible, think about all the tasks that you have done with the system while you answer these questions.
Please read each statement and indicate how strongly you agree or disagree with the statement by circling a number on the scale. If a statement does not apply to you, circle N/A.
Please write comments to elaborate on your answers.
After you have completed this questionnaire, I'll go over your answers with you to make sure I
understand all of your responses.
Thank you!
Likert scale
1 = strongly agree
2
3
4
5
6
7 = strongly disagree
1. Overall, I am satisfied with how easy it is to use this system.
2. It was simple to use this system.
3. I could effectively complete the tasks and scenarios using this system.
4. I was able to complete the tasks and scenarios quickly using this system.
5. I was able to efficiently complete the tasks and scenarios using this system.
6. I felt comfortable using this system.
7. It was easy to learn to use this system.
8. I believe I could become productive quickly using this system.
9. The system gave error messages that clearly told me how to fix problems.
10. Whenever I made a mistake using the system, I could recover easily and quickly.
11. The information (such as on-line help, on-screen messages and other documentation)
provided with this system was clear.
12. It was easy to find the information I needed.
13. The information provided for the system was easy to understand.
14. The information was effective in helping me complete the tasks and scenarios.
15. The organization of information on the system screens was clear.
Note: The interface includes those items that you use to interact with the system. For example, some components of the interface are the keyboard, the mouse, the screens (including their use of graphics and language).
16. The interface of this system was pleasant.
17. I liked using the interface of this system.
18. This system has all the functions and capabilities I expect it to have.
19. Overall, I am satisfied with this system.
Appendix Table 1. Rules for Calculating CSUQ/PSSUQ Scores
_____________________________________________________________________________
Score Name > Average the Responses to:
_____________________________________________________________________________
OVERALL > Items 1 through 19
SYSUSE > Items 1 through 8
INFOQUAL > Items 9 through 15
INTERQUAL > Items 16 through 18
_____________________________________________________________________________
IBM Computer Usability Satisfaction Questionnaires:
Psychometric Evaluation and Instructions for Use
Technical Report 54.786
James R. Lewis
Human Factors Group
Boca Raton, FL
Source: http://drjim.0catch.com/usabqtr.pdf
Administration and Scoring.
Give the PSSUQ to participants after they have completed all the scenarios in a usability study.
You can calculate four scores from the responses to the PSSUQ items:
* the overall satisfaction score (OVERALL),
* system usefulness (SYSUSE),
* information quality (INFOQUAL) and
* interface quality (INTERQUAL).
Because research on an alternative form of the PSSUQ (the Computer System Usability Questionnaire, or CSUQ) confirmed and clarified (and slightly modified) the factor structure of the questionnaire, refer to Appendix Table 1 in the next section of this appendix for the current scoring rules of the PSSUQ.
Instructions and Items.
The questionnaire's instructions and items are:
This questionnaire, which starts on the following page, gives you an opportunity to tell us your reactions to the system you used. Your responses will help us understand what aspects of the system you are particularly concerned about and the aspects that satisfy you.
To as great a degree as possible, think about all the tasks that you have done with the system while you answer these questions.
Please read each statement and indicate how strongly you agree or disagree with the statement by circling a number on the scale. If a statement does not apply to you, circle N/A.
Please write comments to elaborate on your answers.
After you have completed this questionnaire, I'll go over your answers with you to make sure I
understand all of your responses.
Thank you!
Likert scale
1 = strongly agree
2
3
4
5
6
7 = strongly disagree
1. Overall, I am satisfied with how easy it is to use this system.
2. It was simple to use this system.
3. I could effectively complete the tasks and scenarios using this system.
4. I was able to complete the tasks and scenarios quickly using this system.
5. I was able to efficiently complete the tasks and scenarios using this system.
6. I felt comfortable using this system.
7. It was easy to learn to use this system.
8. I believe I could become productive quickly using this system.
9. The system gave error messages that clearly told me how to fix problems.
10. Whenever I made a mistake using the system, I could recover easily and quickly.
11. The information (such as on-line help, on-screen messages and other documentation)
provided with this system was clear.
12. It was easy to find the information I needed.
13. The information provided for the system was easy to understand.
14. The information was effective in helping me complete the tasks and scenarios.
15. The organization of information on the system screens was clear.
Note: The interface includes those items that you use to interact with the system. For example, some components of the interface are the keyboard, the mouse, the screens (including their use of graphics and language).
16. The interface of this system was pleasant.
17. I liked using the interface of this system.
18. This system has all the functions and capabilities I expect it to have.
19. Overall, I am satisfied with this system.
Appendix Table 1. Rules for Calculating CSUQ/PSSUQ Scores
_____________________________________________________________________________
Score Name > Average the Responses to:
_____________________________________________________________________________
OVERALL > Items 1 through 19
SYSUSE > Items 1 through 8
INFOQUAL > Items 9 through 15
INTERQUAL > Items 16 through 18
_____________________________________________________________________________
IBM Computer Usability Satisfaction Questionnaires:
Psychometric Evaluation and Instructions for Use
Technical Report 54.786
James R. Lewis
Human Factors Group
Boca Raton, FL
Source: http://drjim.0catch.com/usabqtr.pdf
Oct 9 - HCIRN, Usability Questionnaires
Usability questionnaires intend to assess users' perception of the usability of a product.
A number of usability questionnaires have been developed or are currently under development:
ASQ (After-Scenario Questionnaire)
CSUQ (Computer System Usability Questionnaire)
CUSI (Computer User Satisfaction Inventory)
ErgoNorm Questionnaire
IsoMetrics
ISONORM 9241/10
MUMMS (Measuring the Usability of Multi-Media Systems)
PSSUQ (Post-Study System Usability Questionnaire)
PUEU (Perceived Usefulness and Ease of Use)
PUTQ (Purdue Usability Testing Questionnaire)
QUIS (Questionnaire for User Interaction Satisfaction)
SUMI (Software Usability Measurement Inventory)
SUS (System Usability Scale)
USE (Usefulness, Satisfaction, and Ease of Use)
WAMMI (Website Analysis and MeasureMent Inventory)
Some of these questionnaires are free, others are commercial and require a license to use them.
Two popular commercial questionnaires are QUIS (Questionnaire for User Interaction Satisfaction) and SUMI (Software Usability Measurement Inventory). Both provided analysis software with output in various formats.
Two popular free questionnaires are PSSUQ (Post-Study System Usability Questionnaire) and SUS (System Usability Scale).
Two of the questionnaires in the above list are related to the seven dialog principles described in ISO 9241 Part 10. IsoMetrics is directly based on these dialog principles and intends to measure the usability of a system along these dimensions. SUMI has five factors, at least four of which seem to directly correspond to dialog principles.
Global vs. Analytic Evaluation
Usability questionnaires generally provide only global measures of system usability. They do not provide any information about which aspects of the system contribute to the positive or negative evaluation. (Many authors use the terms formative and summative evaluation to distinguish between these types evaluation. However, based on a detailed analysis of these terms we believe that the terms analytic and global evaluation are more appropriate. See Formative and Summative Evaluation for more detail.)
Providing only a global measure is a major limitation of usability questionnaires. A number of questionnaires, therefore, attempt to facilitate analytic evaluation as well. One common way which nearly all usability questionnaires support is to provide a comments sections in which respondents can enter feedback in free form.
Some questionnaires also support some form of item analysis. A number of case studies using QUIS (Questionnaire for User Interaction Satisfaction) compared average item scores against the overall average score. However, this approach is questionable.
SUMI (Software Usability Measurement Inventory) uses a method called Item Consensual Analysis (ICA) to obtain additional information about usability problems. Item Consensual Analysis compares the response pattern of each item with its response pattern in the standardization database.
Assessing a Single System vs. Comparing Two or More Systems
Questionnaires can be used to assess a single system or to compare two or more systems. When using a questionnaire to assess a single system, the resulting score by itself is not very meaningful. What does it mean that your system has a score of 5.0 on a range from 1.0 to 7.0? It is better or worse than other comparable systems? To interpret a single score, you need a baseline to compare it with.
However, for a baseline to be useful, it has to be updated regularly. If the baseline score was 5.0 two years ago, it could be quite different now. Improvements to applications could have increased the baseline score to 6.0. Conversely, changes to applications may have increased their complexity which decreased the baseline score to 4.0.
Also, baselines need to be specific for the type of application. For example, new products, such as web browsers, may initially have a poorer score than established products, such as word processors. So a score of 5.0 might be poor compared to word processors and excellent compared to web browsers.
Finally, different baselines are needed for different countries. There might be cultural differences in the interpretation of specific questions as well as in the response tendency of the respondents. This makes it impossible to compare a score to a baseline obtained for another language version of the same questionnaire.
To our knowledge, to date only SUMI (Software Usability Measurement Inventory) provides such a standardization database.
Baselines are not required when comparing two or more systems. All usability questionnaires, provided they have sufficient reliability, are suitable for this. They can be used to measure changes in usability between versions of the same system or to assess differences between competing products.
Validity of Usability Questionnaires
Usability questionnaires can be interpreted in two ways. Firstly, they are a direct measure of users' perception of the usability of a system. Secondly, they are an indirect measure of the actual usability of a system.
...............
...............
One way to validate a questionnaire is to compare it to other questionnaires that claim to measure the same construct. However, few such studies have been conducted for usability questionnaires. Kirakowski (1994) reports two studies that directly compared CUSI (Computer User Satisfaction Inventory), SUS (System Usability Scale) and QUIS (Questionnaire for User Interaction Satisfaction) 5.0. Lucey (1991) found good correlations between the affect subscale of CUSI with SUS and overall QUIS. The correlations between the competence subscale of CUSI with SUS and QUIS were low. Wong & Rengger (1990) found a similar pattern.
Choosing a Usability Questionnaire
Which usability questionnaire should you choose? We have not had the opportunity to examine SUMI (Software Usability Measurement Inventory).
SUMI has never been published and is only available by purchasing the questionnaire package. However, it looks to be your best choice. It is a well-designed and extensively tested questionnaire. It is available in various languages and is the only usability questionnaire that has been standardized. Besides producing scores for a global scale and five subscales, detailed information about a system can be obtained by analyzing the response pattern of each item. The only disadvantage of SUMI is that it costs money. It is worth it, but HCI practitioners might find it difficult to convince management that the expense is justified.
Out of the freely available questionnaires, SUS (System Usability Scale) is probably your best choice. It does not have the standardization database of SUMI and it provides only an overall usability score. However, it is short and less susceptible to response bias than some other freely available questionnaires. SUS is an excellent instrument for comparing systems.
Compared to other evaluation methods in HCI, there is little literature available on usability questionnaires. There are also a number of weak spots in the published research on usability questionnaires. In particular, extensive validation studies that compare usability questionnaires with each other and with other usability evaluation methods are missing. The reason for this is not clear. Perhaps this is an indication that usability questionnaires are currently of little relevance in HCI practice.
Kirakowski, 1994 contains a good introduction to usability questionnaires.
Links
Questionnaires in Usability Engineering: A List of Frequently Asked Questions
By Jurek Kirakowski
http://www.ucc.ie/hfrg/resources/qfaq1.html
UsabilityNet > Questionnaire Resources
http://www.usabilitynet.org/tools/r_questionnaire.htm
A short overview of usability questionnaires.
Web-Based User Interface Evaluation with Questionnaires
By Gary Perlman
http://www.acm.org/~perlman/question.html
Includes online versions of a number of usability questionnaires.
Source: http://www.hcirn.com/atoz/atozu/usaques.php
A number of usability questionnaires have been developed or are currently under development:
ASQ (After-Scenario Questionnaire)
CSUQ (Computer System Usability Questionnaire)
CUSI (Computer User Satisfaction Inventory)
ErgoNorm Questionnaire
IsoMetrics
ISONORM 9241/10
MUMMS (Measuring the Usability of Multi-Media Systems)
PSSUQ (Post-Study System Usability Questionnaire)
PUEU (Perceived Usefulness and Ease of Use)
PUTQ (Purdue Usability Testing Questionnaire)
QUIS (Questionnaire for User Interaction Satisfaction)
SUMI (Software Usability Measurement Inventory)
SUS (System Usability Scale)
USE (Usefulness, Satisfaction, and Ease of Use)
WAMMI (Website Analysis and MeasureMent Inventory)
Some of these questionnaires are free, others are commercial and require a license to use them.
Two popular commercial questionnaires are QUIS (Questionnaire for User Interaction Satisfaction) and SUMI (Software Usability Measurement Inventory). Both provided analysis software with output in various formats.
Two popular free questionnaires are PSSUQ (Post-Study System Usability Questionnaire) and SUS (System Usability Scale).
Two of the questionnaires in the above list are related to the seven dialog principles described in ISO 9241 Part 10. IsoMetrics is directly based on these dialog principles and intends to measure the usability of a system along these dimensions. SUMI has five factors, at least four of which seem to directly correspond to dialog principles.
Global vs. Analytic Evaluation
Usability questionnaires generally provide only global measures of system usability. They do not provide any information about which aspects of the system contribute to the positive or negative evaluation. (Many authors use the terms formative and summative evaluation to distinguish between these types evaluation. However, based on a detailed analysis of these terms we believe that the terms analytic and global evaluation are more appropriate. See Formative and Summative Evaluation for more detail.)
Providing only a global measure is a major limitation of usability questionnaires. A number of questionnaires, therefore, attempt to facilitate analytic evaluation as well. One common way which nearly all usability questionnaires support is to provide a comments sections in which respondents can enter feedback in free form.
Some questionnaires also support some form of item analysis. A number of case studies using QUIS (Questionnaire for User Interaction Satisfaction) compared average item scores against the overall average score. However, this approach is questionable.
SUMI (Software Usability Measurement Inventory) uses a method called Item Consensual Analysis (ICA) to obtain additional information about usability problems. Item Consensual Analysis compares the response pattern of each item with its response pattern in the standardization database.
Assessing a Single System vs. Comparing Two or More Systems
Questionnaires can be used to assess a single system or to compare two or more systems. When using a questionnaire to assess a single system, the resulting score by itself is not very meaningful. What does it mean that your system has a score of 5.0 on a range from 1.0 to 7.0? It is better or worse than other comparable systems? To interpret a single score, you need a baseline to compare it with.
However, for a baseline to be useful, it has to be updated regularly. If the baseline score was 5.0 two years ago, it could be quite different now. Improvements to applications could have increased the baseline score to 6.0. Conversely, changes to applications may have increased their complexity which decreased the baseline score to 4.0.
Also, baselines need to be specific for the type of application. For example, new products, such as web browsers, may initially have a poorer score than established products, such as word processors. So a score of 5.0 might be poor compared to word processors and excellent compared to web browsers.
Finally, different baselines are needed for different countries. There might be cultural differences in the interpretation of specific questions as well as in the response tendency of the respondents. This makes it impossible to compare a score to a baseline obtained for another language version of the same questionnaire.
To our knowledge, to date only SUMI (Software Usability Measurement Inventory) provides such a standardization database.
Baselines are not required when comparing two or more systems. All usability questionnaires, provided they have sufficient reliability, are suitable for this. They can be used to measure changes in usability between versions of the same system or to assess differences between competing products.
Validity of Usability Questionnaires
Usability questionnaires can be interpreted in two ways. Firstly, they are a direct measure of users' perception of the usability of a system. Secondly, they are an indirect measure of the actual usability of a system.
...............
...............
One way to validate a questionnaire is to compare it to other questionnaires that claim to measure the same construct. However, few such studies have been conducted for usability questionnaires. Kirakowski (1994) reports two studies that directly compared CUSI (Computer User Satisfaction Inventory), SUS (System Usability Scale) and QUIS (Questionnaire for User Interaction Satisfaction) 5.0. Lucey (1991) found good correlations between the affect subscale of CUSI with SUS and overall QUIS. The correlations between the competence subscale of CUSI with SUS and QUIS were low. Wong & Rengger (1990) found a similar pattern.
Choosing a Usability Questionnaire
Which usability questionnaire should you choose? We have not had the opportunity to examine SUMI (Software Usability Measurement Inventory).
SUMI has never been published and is only available by purchasing the questionnaire package. However, it looks to be your best choice. It is a well-designed and extensively tested questionnaire. It is available in various languages and is the only usability questionnaire that has been standardized. Besides producing scores for a global scale and five subscales, detailed information about a system can be obtained by analyzing the response pattern of each item. The only disadvantage of SUMI is that it costs money. It is worth it, but HCI practitioners might find it difficult to convince management that the expense is justified.
Out of the freely available questionnaires, SUS (System Usability Scale) is probably your best choice. It does not have the standardization database of SUMI and it provides only an overall usability score. However, it is short and less susceptible to response bias than some other freely available questionnaires. SUS is an excellent instrument for comparing systems.
Compared to other evaluation methods in HCI, there is little literature available on usability questionnaires. There are also a number of weak spots in the published research on usability questionnaires. In particular, extensive validation studies that compare usability questionnaires with each other and with other usability evaluation methods are missing. The reason for this is not clear. Perhaps this is an indication that usability questionnaires are currently of little relevance in HCI practice.
Kirakowski, 1994 contains a good introduction to usability questionnaires.
Links
Questionnaires in Usability Engineering: A List of Frequently Asked Questions
By Jurek Kirakowski
http://www.ucc.ie/hfrg/resources/qfaq1.html
UsabilityNet > Questionnaire Resources
http://www.usabilitynet.org/tools/r_questionnaire.htm
A short overview of usability questionnaires.
Web-Based User Interface Evaluation with Questionnaires
By Gary Perlman
http://www.acm.org/~perlman/question.html
Includes online versions of a number of usability questionnaires.
Source: http://www.hcirn.com/atoz/atozu/usaques.php
Friday, September 25, 2009
Plan for week Sep 28 - Oct 3: Usability Questionnaire
For next week, Sep 28 till Oct 3, I plan to read on Usability Questionnaire.
Usability Questionnaire:
SUMI
QUIS
PSSUQ
SUS
ASQ
CUSI
Usability Questionnaire is also a method used for evaluating usability. In my readings, I have come accross this methodology. I know the Usability Questionnaire would contain a set of questions.
I would like to know more about this. Usability Questionnaire is a type of usability evaluation tool. My research title is "Usability Evaluation Tool for Mobile Learning Applications."
Hence, reading up (literature review) on the popular types of Usability Questionnaire would definitely be useful for my knowledge.
Usability Questionnaire:
SUMI
QUIS
PSSUQ
SUS
ASQ
CUSI
Usability Questionnaire is also a method used for evaluating usability. In my readings, I have come accross this methodology. I know the Usability Questionnaire would contain a set of questions.
I would like to know more about this. Usability Questionnaire is a type of usability evaluation tool. My research title is "Usability Evaluation Tool for Mobile Learning Applications."
Hence, reading up (literature review) on the popular types of Usability Questionnaire would definitely be useful for my knowledge.
Labels:
ASQ,
CUSI,
PSSUQ,
QUIS,
research planning,
SUMI,
SUS,
usability questionnaire
Subscribe to:
Posts (Atom)