Exp1
- 일단 기본적으로 NN만 뽑아서 사용함.(8028개의 단어.)
Corpora Correlation
- 전체 word 간 correlation 비교.
- 구간별 3개로(low, mid, high) 쪼개서 했는데, 각각 KF 기준, (⇐10, 35~50, 100>=) 빈도에 대해서 수행.
- KF 기준으로 한 건, 이 당시에는 대부분들이 이렇게 했다고 만 기술되어 있음.
...
There were four comparisons. The first comparison was over the entire frequency range. The other three comparisons were for low (⇐10), medium (35-75), and high (>=100) frequency ranges, based on frequencies from KF. Unless otherwise specified (i.e., Experiment 2B), WF ranges are based on KF, since this is how most investigators have operationalized WE These operational definitions of LF, medium frequency (MF) and HF were based on our review of the literature of the common usage of these WF ranges in empirical studies done with the KF frequency counts.
Results1
Title
The overall correlation between the two corpora was r= .96. However, when WF was separated into frequency-based categories, there was a lack of correspondence of the frequency estimates from the two corpora. Although the two corpora correlated well for HF words (r = .96, p < .0001; 489 words), the correlations for MF and LF words were considerable weaker (MF, r = .12, P < .000 I, 713 words; LF, r = .14, P < .000 I, 3,720 words). The lack of correspondence between the frequency estimates from the two corpora for the LF and MF words motivated Experiment 2.
전체간 correlation :.96
그러나, 구간별 상관을 구하면 매우 많이 상관 값이 낮아짐.
Exp2
- exp1에서 두 코퍼스가 동일 단어에 대한 다른 크기의 WF를 추정했으므로 행동 값인 RT를 얼마나 잘 예측하는지 비교.

- 실험에서 사용된 단어들은 각각 Brown에서 추출되기도 하고, HAL에서 추출되기도 함.
- 처음에는 exp1 설명에서처럼 KF 코퍼스 기준으로 고,중,저 빈도 categorize하고 이것만 사용하려고 했으나, KF에서 low인게 HAL에서는 mid 로 잡히기도 하는 경우가 발생하여 HAL 기준으로 분석할 때에는 단어의 group label을 번경함. 즉, 단어 리스트는 유지하고 단어의 group만 바꿔서 재 분석.
Results & Discussion
- 규모가 큰 HAL이 전 항목에서 압승.
- HAL corpus 기반의 word-categorization은 두 코퍼스 간 설명력을 더 크게 만들었는데, 왜지?
KF 와 HAL word-freq를 사용해서 RT predictability를 검증.