<?xml version="1.0" encoding="utf-8"?><!DOCTYPE article  PUBLIC '-//OASIS//DTD DocBook XML V4.4//EN'  'http://www.docbook.org/xml/4.4/docbookx.dtd'><article><articleinfo><title>NorasData</title><revhistory><revision><revnumber>2</revnumber><date>2013-08-02 16:04:34</date><authorinitials>RussellThompson</authorinitials></revision></revhistory></articleinfo><para>Open R </para><para>Load all the packages that are needed </para><para>Import the data into R’<!--RAW HTML: &#8217;-->s workspace with </para><screen><![CDATA[n = read.csv("/Users/mcolthea/Desktop/RFolder/Nora/NoraData.csv")]]></screen><para>Now the file called n is the whole raw data file </para><para>Check that it looks OK by reading the first 10 rows of the file: </para><screen><![CDATA[> head(n, n = 10)]]></screen><para>Looks OK. </para><para>Now, we aren’<!--RAW HTML: &#8217;-->t going to want to include errors when we analyse RTs. So the first thing we want to do is create a file that only contains the correct responses </para><screen><![CDATA[> nc = n[n$accuracy=="correct",]]]></screen><para>This new file nc is the file that just has the correct responses. </para><para>What was the overall error rate? </para><screen><![CDATA[> (nrow(n)-nrow(nc))/nrow(n)
[1] 0.0519774]]></screen><para>5.2% of responses were errors. </para><para>One other thing is typically done before data analysis starts: eliminating any responses where the RT is suspiciously long or suspiciously short, i.e. “<!--RAW HTML: &#8220;-->outliers”<!--RAW HTML: &#8221;-->. This is often called “<!--RAW HTML: &#8220;-->data-trimming”<!--RAW HTML: &#8221;-->. Here you are eliminating noise from the data and so making subsequent statistical analyses more sensitive. </para><para>You can visualize the outliers in your data by inspecting the quantile-quantile plot for each subject. Here’<!--RAW HTML: &#8217;-->s how you do this in R: </para><para>First, set things up so that you won’<!--RAW HTML: &#8217;-->t have too many plots on each page; to get to the next page, you press Return </para><screen><![CDATA[> devAskNewPage(ask=TRUE)]]></screen><para>Now display the quantile-quantile plots </para><screen><![CDATA[> qqmath(~RT|Subject.ID, data = nc, layout=c(3,3,4))]]></screen><para>This says there are 3 rows and 3 columns per page, and 4 pages. </para><para>Nice clean data. Not many outliers. But there are a few obvious ones. </para><para>The best way to look at these is to sort each subject’<!--RAW HTML: &#8217;-->s RTs from lowest to highest, and look at these sorted files. They can be sorted in R with this command </para><screen><![CDATA[> ncs=nc[order(nc$Subject.ID, nc$RT),]]]></screen><para>which sorts by Subject.ID and within each subject sorts by RT. </para><para>Trimming the data is more easily done in Excel, so export this sorted trial. </para><screen><![CDATA[>write.csv(ncs,"/Users/mcolthea/Desktop/RFolder/Nora/ncs.csv&#8221;)]]></screen><para>and then open it in Excel. Do the trimming in Excel. Call the trimmed file ncst. Import it to R. </para><screen><![CDATA[> ncst = read.csv("/Users/mcolthea/Desktop/RFolder/Nora/ncst.csv")]]></screen><para>What is the proportion of correct responses that have been discarded? </para><screen><![CDATA[> (nrow(ncs)-nrow(ncst))/nrow(ncs)
[1] 0.01102503]]></screen><para>1.1% of the data have been trimmed. </para><para>Look at the quantile-quantile plots again. </para><screen><![CDATA[> devAskNewPage(ask=TRUE)
> qqmath(~RT|Subject.ID, data = ncst, layout=c(3,3,4))]]></screen><para>The obvious outliers are gone. Now data analysis can begin using this trimmed data. </para><section><title>Analysing the data</title><para>What is the mean RT for each condition? </para><screen><![CDATA[> tapply(ncst$RT, ncst$itemtype, mean)
]]><![CDATA[
   count     mass 
514.0556 523.8704]]></screen><para>Is this mean difference of 9.8 msec significant? </para><screen><![CDATA[> ncst.lmer=lmer(RT~itemtype + (1|Subject.ID) + (1|itemtype), data = ncst )]]></screen><para>You can see the result of the analysis by typing: </para><screen><![CDATA[>ncst.lmer
]]><![CDATA[
and if you do, you get
]]><![CDATA[
Linear mixed model fit by REML 
Formula: RT ~ itemtype + (1 | Subject.ID) + (1 | Item
   Data: ncst 
   
  AIC   BIC  logLik   deviance REMLdev
 39988 40018 -19989    39990   39978
]]><![CDATA[
Random effects:
 Groups     Name        Variance Std.Dev.
 Subject.ID (Intercept) 3262.700 57.1201 
 itemtype   (Intercept)   14.890  3.8588 
 Residual               9682.982 98.4021 
Number of obs: 3319, groups: Subject.ID, 30; itemtype, 2
]]><![CDATA[
Fixed effects:
             Estimate Std. Error t value
(Intercept)   513.274     11.295   45.44
itemtypemass   10.128      6.297    1.61
]]><![CDATA[
Correlation of Fixed Effects:
            (Intr)
itemtypemss -0.265]]></screen><para>What does this all mean? </para><itemizedlist><listitem override="none"><para>The t-value is 1.61 which wouldn’<!--RAW HTML: &#8217;-->t be significant, but even if it were, we haven’<!--RAW HTML: &#8217;-->t taken into account the potential confounding factor of  LOG_Surface.freq. So add that factor to the model </para></listitem></itemizedlist><screen><![CDATA[> ncst.lmerA=lmer(RT~itemtype + (1|Subject.ID) + (1|Item) +LOG_Surface.freq, data = ncst )]]></screen><para>And what about LOG_stem.freq? </para><screen><![CDATA[> ncst.lmerB=lmer(RT~itemtype + (1|Subject.ID) + (1|Item) +LOG_Surface.freq+LOG_stem.freq , data = ncst )]]></screen><para>So it looks as if there is no mass/count effect nor a surface frequency effect, but there is a stem frequency effect. Since we have a subset of the items that are matched on stem frequency, analyzing that subset next would be a good move. </para><para>Another thing we can do is build a set of different models incorporating different factors. </para><screen><![CDATA[> ncst.lmerC=lmer(RT~(1|Subject.ID) + (1|Item) + LOG_stem.freq , data = ncst )]]></screen><para>includes only the stem frequency factor. </para><screen><![CDATA[>ncst.lmerD=lmer(RT~(1|Subject.ID) + (1|Item) + LOG_stem.freq +itemtype , data = ncst )]]></screen><para>adds a second factor, item type. Does the second model fit better than the first? </para><screen><![CDATA[> anova (ncst.lmerC,ncst.lmerD)
]]><![CDATA[
Data: ncst
Models:
ncst.lmerC: RT ~ (1 | Subject.ID) + (1 | Item) + LOG_stem.freq
ncst.lmerD: RT ~ (1 | Subject.ID) + (1 | Item) + LOG_stem.freq + itemtype
           Df   AIC   BIC logLik  Chisq Chi Df Pr(>Chisq)
ncst.lmerC  5 39802 39832 -19896                         
ncst.lmerD  6 39804 39840 -19896 0.3083    1.     0.5787]]></screen><para>No. The insignificant Chi Squared shows that a model that includes item type and stem frequency does not produce a better fit than a model that only includes stem frequency. </para><para>Here’<!--RAW HTML: &#8217;-->s a model that has surface frequency and item type: </para><screen><![CDATA[> ncst.lmerG=lmer(RT~(1|Subject.ID) + (1|Item) +LOG_Surface.freq +itemtype  , data = ncst )]]></screen><para>If we add stem frequency to this model </para><screen><![CDATA[> ncst.lmerF=lmer(RT~(1|Subject.ID) + (1|Item) + LOG_stem.freq +LOG_Surface.freq +itemtype  , data = ncst )]]></screen><para>Do we get a better fit? Does model F perform better than model G? </para><screen><![CDATA[> anova (ncst.lmerG,ncst.lmerF)
Data: ncst
Models:
ncst.lmerG: RT ~ (1 | Subject.ID) + (1 | Item+ LOG_Surface.freq + itemtype
ncst.lmerF: RT ~ (1 | Subject.ID) + (1 | Item+ LOG_stem.freq + LOG_Surface.freq + 
ncst.lmerF:     itemtype
           Df   AIC   BIC logLik  Chisq Chi Df Pr(>Chisq)    
ncst.lmerG  6 39817 39854 -19903                             
ncst.lmerF  7 39805 39848 -19896 14.448    1.  0.0001441 ***]]></screen><para>Yes. G is a significant improvement over F. Note how the various measures of fit are lower for F. That means the fit is closer. </para><para>So adding stem frequency gives a s a better model, but adding item type or surface frequency or both does not improve the fit of the model to the data. </para></section></article>