Showing posts with label CSUQ. Show all posts
Showing posts with label CSUQ. Show all posts

Friday, October 9, 2009

Oct 10 - Tullis & Stetson, A Comparison of Questionnaires for Assessing Website Usability

A Comparison of Questionnaires for Assessing Website Usability
Thomas S. Tullis, Fidelity Investments
Jacqueline N. Stetson, Fidelity Investments and Bentley College

 
Various questionnaires have been reported in the literature for assessing the perceived usability of an interactive system, e.g:
–Questionnaire for User Interface Satisfaction (QUIS) (1988)
–Computer System Usability Questionnaire (CSUQ) (1995)
–System Usability Scale (SUS) (1996)
 
A slightly different approach was taken by Microsoft with their "Product Reaction Cards"(2002)
And we have been using our own questionnaire for several years in our Usability Lab at Fidelity Investments
 
Problem
How well do these questionnaires apply to the assessment of Websites?
Do any of these questionnaires work well, as an adjunct to a usability test, with relatively small numbers of users?


Our Study
Limited ourselves to questionnaires in the published literature
–Did not include commercial services for evaluating website usability (e.g., WAMMI, RelevantView, NetRaker, Vividence).
We studied five questionnaires:
–SUS
–QUIS
–CSUQ
–Microsoft’s "Words"
–Our own questionnaire
 
Questionnaire #1: SUS
•Developed at Digital Equipment Corp.
•Consists of ten items.
•Adapted by replacing "system"with "website".
•Each item is a statement (positive or negative) and a rating on a five-point scale of "Strongly Disagree"to "Strongly Agree".
 
Questionnaire #2: QUIS
•Developed at the University of Maryland.
•Original questionnaire had 27 questions.
–We dropped 3 that did not seem relevant to Websites (e.g., "Remembering names and use of commands").
•"System"was replaced by "website"and term "screen"was replaced by "web page".
•Each question is a rating on a ten-point scale with appropriate anchors.
 
Questionnaire #3: CSUQ
•Developed at IBM.
•Composed of 19 questions.
•"System"or "computer system"was replaced by "website".
•Each question is a statement and a rating on a seven-point scale of "Strongly Disagree"to "Strongly Agree".
 
Questionnaire #4: Words
•Based on the 118 words used by Microsoft on their Product Reaction Cards.
–Some positive (e.g., "Convenient")
–Some negative (e.g., "Unattractive")
•Each word was presented with a check-box
–Users were asked to choose the words that best describe their interaction with the website.
–Could choose as many or as few words as they wished.

Questionnaire #5: Ours
•Developed ourselves and have been using for several years in our usability tests of websites.
•Composed of nine statements (e.g., "This website is visually appealing") to which the user responds on a seven-point scale from "Strongly Disagree"to "Strongly Agree".
•Points of the scale are numbered -3, -2, -1, 0, 1, 2, 3.
–Obvious neutral point at 0.
 
A Live Experiment!
•We’re going to compare two sites:
–CircuitCity.com
–Outpost.com
•Task 1: Your digital camera uses SmartMediacards. Find the least expensive external reader (USB) for your PC that will read them.
•Task 2: You do lots of hiking. Find the least
expensive personal GPS with map capability
and at least 8 MB of memory.

 
Method of Our Study
•Conducted entirely on our company Intranet.
•123 of our employees participated.
•Each participant was randomly assigned to one of the five questionnaire conditions.
•Each was asked to perform two tasks on each of two well-known personal financial information sites.
•Sites studied:
–Finance.Yahoo.com
–Kiplinger.com
–Hereafter referred to only as "Site 1"and
"Site 2". Don’t assume which is which.
•Tasks:
–Find the highest price in the past year for a share of .
–Find the mutual fund with the highest 3-year return.
•Order of presentation of the two sites was randomized.
•After completing (or at least attempting) the two tasks on a site, the user was presented with the questionnaire for their randomly selected condition.
•Each user completed the same questionnaire for both sites.

 
Data Analysis
•For each participant, an overall score was calculated for each website by averaging all of the ratings on the questionnaire that was used.
–All scales had been coded internally so that the "better"end corresponded to higher numbers.
–These were converted to percentages by dividing each score by the maximum score possible on that scale.
–For example, a rating of 3 on SUS was converted to a percentage by dividing that by 5 (the maximum score for SUS), giving a percentage of 60%.
•Special treatment for the "Words"condition since it did not involve rating scales:
–Before the study, we classified each of the words as being "Positive"or "Negative".
–Not grouped or identified as such to the participants.
–For each participant, an overall score was calculated by counting the total number of words that person selected and then dividing that number into the number of "Positive"words chosen.
–If someone selected 8 positive words and 10 words total, that yielded a score of 80%.
 
Results
•Calculated frequency distributions for the ratings, converted to percentages, for:
–Each questionnaire
–Both websites

Bar Charts are used to compare the results. See http://www.upassoc.org/usability_resources/conference/2004/UPA-2004-TullisStetson.pdf

My Comments: This study will be a good benchmark and reference for my research.

 
Results: Summary
•All five questionnaires showed that Site 1 was significantly preferred over Site 2 (p<.01). •The largest mean difference (74% vs. 38%) was found using the Words questionnaire, but this was also the questionnaire that yielded the greatest variability. Analysis of Sub-samples
•Next we analyzed randomly selected sub-samples of the data at size 6, 8, 10, 12, and 14.
–20 random samples for each size
•For each sample, t-test was conducted to determine whether the results showed that Site 1 was significantly better than Site 2 (the conclusion from the full dataset).
 
Analysis of Sub-samples
•Accuracy of the results increases as the sample size gets larger.
•With a sample size of only 6, all of the questionnaires yield accuracy of only 30-40%
–60-70% of the time, at that sample size, you would fail to find a significant difference between the two sites.
•Accuracy of some of the questionnaires increases quicker than others.
–SUS jumps up to about 75% accuracy at a size of 8.
 
Caveats
•Results were undoubtedly influenced by:
–The sites studied.
–The tasks used.
•We have only addressed the question of whether a given questionnaire was able to reliably distinguish between the ratings of one site vs. the other.
–Often you care more about how well the results help guide a redesign.
 

Conclusions
•One of the simplest questionnaires studied, SUS (with only 10 rating scales), yielded among the most reliable results across sample sizes.
–Also the only one whose questions all address different aspects of the user’s reaction to the website as a whole.
•For the conditions of this study, sample sizes of at least 12-14 participants are needed to get reasonably reliable results.

Source:
http://www.upassoc.org/usability_resources/conference/2004/UPA-2004-TullisStetson.pdf

Oct 10 - Usabilitynet, Usability Questionnaires

My comments: This is a good summary/intro of several Usability Questionnaires.

It has been observed that questionnaires are the most frequently used tools for usability evaluation. This page is a list of usability questionnaire resources, extending the information presented on the questionnaires page of Usabilitynet.

SUMI
This is a mature questionnaire whose standardisation base and manual have been regularly updated. It is good for desktop products, but has also been used to evaluate command-and-control applications. It is a commercial product which comes complete with scoring and report generation software. It is designed and sold by the Human Factors Research Group at University College Cork.

WAMMI
This is a new questionnaire, designed to evaluate the quality of use of web sites. It is backed up by an extensive standardisation database, and it is purchased on a per report basis. It is the result of a joint development project by Jurek Kirakowski and Nigel Claridge.

SUS
This is a mature questionnaire, developed by John Brooke in 1986 and not published until years later. It is very robust and has been extensively used and adapted. It is public domain and nobody has published any standardisation data about it. Of all the public domain questionnaires, this is the most strongly recommended.

QUIS
This is a questionnaire developed by Kent Norman that has been modified many times to keep it current since its first appearance. It is commercially available and is championed by Ben Shneiderman in his book Designing the User Interface. Reading, MA: Addison-Wesley Publishing Co., 1998. Despite lack of standardisation and validation data, it has many adherents.

USE (see http://www.mindspring.com/~alund/USE/IntroductionToUse.html)
This questionnaire is still in development by Arnie Lund (last updated 11/11/98). It attempts to create a three-factor model of usability that can be applied to many situations. However, no reliability or validation data are presented. Public domain use is encouraged.

CSUQ
This is a well-designed questionnaire developed by Jim Lewis and it is public domain. It has excellent psychometric reliability properties but no standardisation base.

IsoNorm (in German only)
This questionnaire is designed to test the usability quality of software following the ISO 9241 part 10 principles. It is created by a team led by Jochim Puemper. Strong reliabilities are claimed for the sub-scales, although it appears there may be a strong inter-correlation between them as well. Downloads and an on-line version are available from the above URL, as well as articles about it (all in German.)

IsoMetrics
This questionnaire is produced by Guenter Gediga and his team. It is another attempt to produce a way of measuring ISO 9241 part 10, with reference to specific software features that may give rise to low usability data. It is therefore good both for summative and formative assessments. The questionnaire is well researched and detailed statistical information is given. Downloads of English and German versions are available. There is no standardisation base for it but it is public domain.


Questionnaire resources
Source: http://www.usabilitynet.org/tools/r_questionnaire.htm

Oct 9 - CSUQ: Computer System Usability Questionnaire - Lewis of IBM

The Computer System Usability Questionnaire (CSUQ)

The PSSUQ research was preliminary for two reasons. First, the sample size for the factor analysis was small, consisting of data from only 48 participants. Second, the PSSUQ data came from a usability study. This setting may have influenced the correlations among the items and, therefore, the resultant factors.
The purpose of this research (Lewis, 1992a) was to use a slightly revised version of the PSSUQ, the Computer System Usability Questionnaire (CSUQ) to obtain a database of sufficient size to calculate stable factors from a mailed survey.
If the same factors emerged from this research as from the PSSUQ research, the study would demonstrate the potential usefulness of the questionnaire across different user groups and different research settings.

Item Selection and Construction
The CSUQ is identical to the PSSUQ (Lewis, 1991c), except that the wording of the items does not refer to a usability testing situation. For example, Item 3 of the PSSUQ states, "I could effectively complete the tasks and scenarios using this system," but Item 3 of the CSUQ states, "I can effectively complete my work using this system." (See the appendix for the CSUQ items.)

Psychometric Evaluation

The mail survey using the CSUQ.
The participants were 825 IBM employees who worked at nine IBM development sites: Atlanta, Austin, Bethesda, Boca Raton, Dallas, Raleigh, Rochester, San Jose, and Tucson. I used a random number generator to select the participants' names from the IBM electronic mail directory (CALLUP), and mailed them each a copy of the CSUQ with a cover letter. Responses from the returned questionnaires that arrived within 3 months of mailing made up the database for this study.

Factor analysis.
Forty-six percent (377) of the participants returned the questionnaire.
A principal factor analysis of the returned questionnaires produced the scree plot shown in Figure 3. The scree plot was similar to that found for the PSSUQ, indicating that an appropriate factor analysis should solve for three factors. Table 6 shows the varimax-rotated 3-factor solution. The selection criterion for the factor loadings was 0.5, shown in bold type in the table.
The factor analysis showed that Item 8 ("I believe I became productive quickly using this system"), which was not a part of the original PSSUQ, should be part of Factor 1. Item 15 ("The organization of information on the system screens is clear"), which loaded on two factors in the PSSUQ study, loaded on only Factor 2 in the current study. In the PSSUQ study and in the current study, Item 19 ("Overall, I am satisfied with this system") loaded on both Factors 1 and 3, and is not part of any subscale.
Otherwise, the factor structure of the CSUQ is very similar to that of the PSSUQ, so the CSUQ and PSSUQ subscales have the same names.
The three factors accounted for 98.6% of the variability in the rating data.

Reliability.
In all cases, coefficient alpha exceeded 0.89, indicating acceptable scale reliability. The estimates of coefficient alpha for the CSUQ were .93 for SYSUSE, .91 for INFOQUAL, .89 for INTERQUAL, and .95 for the OVERALL scale. The values of coefficient alpha for the CSUQ scales were within 0.03 of those for the PSSUQ scales.

Validity/Sensitivity.
After establishing scale reliability, the next step in psychometric evaluation is to determine scale validity. However, without a concurrent or predicted measurement, it is impossible to obtain a quantitative measure of validity in the traditional psychometric sense. An indirect way to assess validity is to examine scale sensitivity to variables that should systematically affect the scale. The sensitivity analyses of the PSSUQ (Lewis, 1992b) showed significant effects of user group (business professional with mouse experience, business professional without mouse experience, and secretary/clerk without mouse experience) on the OVERALL, SYSUSE, INFOQUAL, and INTERQUAL scales. The type of computer system the participant used during the study significantly affected the INFOQUAL scale.

A comprehensive listing of the influence of respondent characteristics on the CSUQ scores is outside the scope of this paper. However, the significant findings are similar to those for the PSSUQ. The type of computer that respondents used significantly affected their responses only for the INFOQUAL score (F(5,311)=2.14, p=0.06). The number of years of experience with their computer system affected respondents' scores for OVERALL (F(4,294)=3.12, p=0.02), SYSUSE (F(4,332)=2.05, p=0.09), INFOQUAL (F(4,311)=2.59, p=0.04) and INTERQUAL (F(4,322)=2.47, p=0.04). The respondents' range of experience with computer systems (number of different computer systems that they reported having used) affected scores for OVERALL (F(3,294)=2.77, p=0.04), INFOQUAL (F(3,311)=2.60, p=0.05) and INTERQUAL (F(3,322)=2.14, p=0.10).
These significant findings provide indirect support to the hypothesis that these scales are valid.

Discussion
The key results from this study are
(1) a demonstration of stable factors for the CSUQ (and, by extension, for the PSSUQ) and
(2) evidence that the questionnaire works well in non-laboratory settings.
The CSUQ scales are comparable to the PSSUQ scales, both in terms of reliability and validity (indicated by similarity in the sensitivity analyses).
These findings substantially enhance the usefulness of the CSUQ and PSSUQ to usability practitioners. Researchers who conduct usability studies (either laboratory or non-laboratory) can use this questionnaire to assess user satisfaction with system usability.


My Comments: This would be helpful if I eventually select to develop Usability Questionnaire as usability evaluation tool.

IBM Computer Usability Satisfaction Questionnaires: Psychometric Evaluation and Instructions for Use
Technical Report 54.786
James R. Lewis
Human Factors Group
Boca Raton, FL

Source: http://drjim.0catch.com/usabqtr.pdf

Thursday, October 8, 2009

Oct 9 - Lewis, IBM Computer Usability Satisfaction Questionnaires: Psychometric Evaluation and Instructions for Use

IBM Computer Usability Satisfaction Questionnaires: Psychometric Evaluation and Instructions for Use


ABSTRACT
This paper describes recent research in subjective usability measurement at IBM. The focus of the research was the application of psychometric methods to the development and evaluation of questionnaires that measure user satisfaction with system usability. The primary goals of this paper are to (1) discuss the psychometric characteristics of four IBM questionnaires that measure user satisfaction with computer system usability, and (2) provide the questionnaires, with administration and scoring instructions. Usability practitioners can use these questionnaires with confidence to help them measure users' satisfaction with the usability of computer systems.


Introduction

Customers want usable products, and developers strive to produce them. It follows that an important part of modern product engineering, both hardware and software, must be the measurement of usability.
Measuring usability is particularly difficult because usability is not a unidimensional product or user characteristic, but emerges as a multidimensional characteristic in the context of users performing tasks with a product in a specific environment (Bevan, Kirakowski, & Maissel, 1991; Shackel, 1984).
However, if you are unable to measure usability, how can you judge your product against your competitors', or even your own previous versions of the product?

Subjective and Objective Evaluation

Most usability evaluations gather both subjective and objective quantitative data in the context of realistic scenarios-of-use, as well as descriptions of the problems representative participants have trying to complete the scenarios.
Subjective data are measures of participants' opinions or attitudes concerning their perception of usability.
Objective data are measures of participants' performance (such as scenario completion time and successful scenario completion rate).

Objective usability measures include, but are not limited to, scenario completion time, successful scenario completion rate, and time spent recovering from errors (Whiteside, Bennett, & Holtzblatt, 1988). Subjective usability measures are usually responses to Likert-type questionnaire items that assess user attitude concerning attributes such as system ease-of-use and interface likeability (Alty, 1992).
Most usability evaluators collect both objective and subjective data.

Research Focus

The focus of this research was the application of psychometric methods to the development and evaluation of standard questionnaires to assess subjective usability.
The goal of psychometrics is to establish the quality of psychological measures (Nunnally, 1978). Is a measure reliable in the sense that it is consistent? Given a reliable measure, is it valid (measures the intended attribute)? Finally, is the measure appropriately sensitive to experimental manipulations?
Psychometrics is a well-developed field, but usability researchers have only recently used these methods to develop and evaluate questionnaires to assess usability (Sweeney & Dillon, 1987).
In contrast to other recent computer-user satisfaction questionnaires (Chin, Diehl, & Norman, 1988; Kirakowski & Dillon, 1988; LaLomia & Sidowski, 1990) the IBM questionnaires are specifically for use in the context of scenario-based usability testing (Lewis, 1991a; Lewis, 1991b; Lewis, 1991c; Lewis, 1992b; Lewis, Henry, & Mack, 1990), although additional research has indicated that one may be useful as an instrument for field evaluation (Lewis, 1992a). Usability practitioners can use these questionnaires to enhance their current usability methods. (The four IBM questionnaires appear in the appendix.)
Before describing the psychometric properties of the IBM questionnaires, I will briefly review the relevant elements of psychometric practice. (For a comprehensive discussion of psychometrics, see Nunnally, 1978.)

----------
ASQ
PSQ
PSSUQ
CSUQ
-----------

General Discussion

Although user satisfaction with system usability is only one component of the multifaceted construct of usability (Bevan et al., 1991), it is a very important component in many situations.
It is especially important when a primary design goal is user satisfaction.

This paper has described the psychometric qualities of four questionnaires that assess user satisfaction with system usability: the ASQ, PSQ, PSSUQ and CSUQ.

The ASQ and PSQ are both after-scenario questionnaires, intended for use in a scenario-based usability testing situation. They contain essentially the same items, but the ASQ uses a 7-point scale and the PSQ uses a 5-point scale.
Using data from very different scenario-based usability studies (one a study of software office applications, the other a study of printers), their factor analyses, validity analyses, and sensitivity analyses were virtually identical. Obtaining the same results in different settings with different user groups provides strong evidence that these results are generalizable, and the questionnaires have wide applicability. Because the ASQ has substantially better reliability than the PSQ, usability practitioners should use the ASQ rather than the PSQ as their after-scenario questionnaire.

The PSSUQ and CSUQ are both overall satisfaction questionnaires. The PSSUQ items are appropriate for a usability testing situation, and the CSUQ items are appropriate for a field testing situation. Otherwise, the questionnaires are identical.
The psychometric evaluations of the PSSUQ (using data from a usability study) and the CSUQ (using data from a mail survey) were virtually identical. As with the after-scenario questionnaires, this consistency provides strong evidence of generalizability of results and wide applicability of the questionnaires.

Because these questionnaires have acceptable psychometric properties, usability practitioners can use them with confidence as standardized measurements of satisfaction for usability studies and tests (ASQ, PSSUQ) or field research (CSUQ).
(Practitioners should note that nothing prevents the addition of items to these questionnaires if a particular situation suggests the need. However, using these questionnaires as the foundation for special-purpose questionnaires ensures that practitioners can score the scales and subscales from the questionnaires, maintaining the advantages of standardized measurement.)

Standardized satisfaction measurements offer many advantages to the usability practitioner (Nunnally, 1978). Specifically, standardized measurements provide:

1 Objectivity.
A standardized measurement supports objectivity because it allows usability practitioners to independently verify the measurement statements of other practitioners.

2 Quantification.
Standardized measurements allow practitioners to report results in finer detail than they could using only personal judgment. Standardization also permits practitioners to use powerful methods of mathematics and statistics to better understand their results (Nunnally, 1978).

3 Communication.
It is easier for practitioners to communicate effectively when standardized measures are available. Inadequate efficiency and fidelity of communication in any field is an impediment to progress.

4 Economy.
Developing standardized measures requires a substantial amount of work. However, once developed, they are economical. There is rarely any need to re-evaluate standardized measures.

5 Scientific generalization.
Scientific generalization is at the heart of scientific work. Standardization is essential for assessing the generalization of results.

Conclusion
In conclusion, these questionnaires should be valuable additions to the repertoire of techniques that usability practitioners apply in the design and evaluation of computer systems.



IBM Computer Usability Satisfaction Questionnaires: Psychometric Evaluation and Instructions for Use
Technical Report 54.786
James R. Lewis
Human Factors Group
Boca Raton, FL

Source: http://drjim.0catch.com/usabqtr.pdf

Oct 9 - Computer System Usability Questionnaire (CSUQ)

Computer System Usability Questionnaire (CSUQ)

Administration and Scoring.
Use the CSUQ rather than the PSSUQ when the
usability study is in a non-laboratory setting. Appendix Table 1 contains the rules for
calculating the CSUQ and PSSUQ scores.
_____________________________________________________________________________
Appendix Table 1. Rules for Calculating CSUQ/PSSUQ Scores
_____________________________________________________________________________
Score Name > Average the Responses to:
_____________________________________________________________________________
OVERALL > Items 1 through 19
SYSUSE > Items 1 through 8
INFOQUAL > Items 9 through 15
INTERQUAL > Items 16 through 18
_____________________________________________________________________________
Average the scores from the appropriate items to obtain the scale and subscale
scores. Low scores are better than high scores due to the anchors used in the 7-point
scales. If a participant does not answer an item or marks "N/A," then average the
remaining item scores.

Instructions and Items.
The questionnaire's instructions and items are:
This questionnaire (which starts on the following page) gives you an opportunity to express your satisfaction with the usability of your primary computer system. Your responses will help us understand what aspects of the system you are particularly concerned about and the aspects that satisfy you.
To as great a degree as possible, think about all the tasks that you have done with the system while you answer these questions.
Please read each statement and indicate how strongly you agree or disagree with the statement by circling a number on the scale. If a statement does not apply to you, circle N/A.
Whenever it is appropriate, please write comments to explain your answers.
Thank you!

Likert scale:
1 = strongly agree
2
3
4
5
6
7 = strongly disagree


1. Overall, I am satisfied with how easy it is to use this system.

2. It is simple to use this system.

3. I can effectively complete my work using this system.

4. I am able to complete my work quickly using this system.

5. I am able to efficiently complete my work using this system.

6. I feel comfortable using this system.

7. It was easy to learn to use this system.

8. I believe I became productive quickly using this system.

9. The system gives error messages that clearly tell me how to fix problems.

10. Whenever I make a mistake using the system, I recover easily and quickly.

11. The information (such as on-line help, on-screen messages and other documentation) provided with this system is clear.

12. It is easy to find the information I need.

13. The information provided with the system is easy to understand.

14. The information is effective in helping me complete my work.

15. The organization of information on the system screens is clear.

Note: The interface includes those items that you use to interact with the system. For example, some components of the interface are the keyboard, the mouse, the screens (including their use of graphics and language).

16. The interface of this system is pleasant.

17. I like using the interface of this system.

18. This system has all the functions and capabilities I expect it to have.

19. Overall, I am satisfied with this system.



IBM Computer Usability Satisfaction Questionnaires: Psychometric Evaluation and Instructions for Use
Technical Report 54.786
James R. Lewis
Human Factors Group
Boca Raton, FL

Source: http://drjim.0catch.com/usabqtr.pdf