<div>Maybe the better word for "equivalence" is "adequateness" or "similarity".</div><div><br></div>I believe there are two types of variability (similarity) we are talking about: George and Mike would study similarity at the grammatical level; Pete at the cognitive level. I suppose that every particular level has its drawbacks :( Semantic similarities between subjects provide a fascinating basis. However, some cultures do not have particular things and therefore no word for this subject. Grammars can be very different.<div>


<div><br></div><div>Since languages are very different, it is probably not feasible to find a "universal" frequency list. For this reason, I would simplify the discussion and limit it to the following question: What properties of two nationalities can be considered similar enough to entail a similar list of the most frequent words? The same grammar, realms, etc? In other words, given language A and language B, what properties of both languages (both grammatical and cognitive) influence the list of the most frequent words? I assume European languages can have similar lists of the most frequent languages because they have very similar realms; language grammar can be also similar.</div>


<div><div><br></div><div>Marvelous examples can be Eastern Germany vs. Western Germany (both speaking the same language but having different realms; American English vs. British English). As Georgios said temporality plays a minor role in this discussion. How about geography? The list of the frequent words in the same same country at the both borders is the same?</div>

<div>


<div><span style="font-family:arial, sans-serif;font-size:13px;background-color:rgb(255, 255, 255)"><br></span></div><div><span style="font-family:arial, sans-serif;font-size:13px;background-color:rgb(255, 255, 255)">Alexander</span></div>


<div><br><div class="gmail_quote">2011/10/10 Georgios Mikros <span dir="ltr"><<a href="mailto:gmikros@isll.uoa.gr" target="_blank">gmikros@isll.uoa.gr</a>></span><br><blockquote class="gmail_quote" style="margin:0 0 0 .8ex;border-left:1px #ccc solid;padding-left:1ex">


<div lang="EN-US" link="blue" vlink="purple"><div><p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">Dear Alexander,<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">The 1000 most frequent words of most languages are mainly function words and their frequency distribution can be predicted with reasonable accuracy using the Zipf’s law. In a number of experiments we have conducted in the early ’00 for Modern Greek [1]  we found that 90% of the 1000 most frequent words do not change even when we triple the size of the corpus (from 13Mwords to 33Mwords) and change considerably its topics and genres structure. So we are dealing probably with a lexical core which due to the grammatical character of its constituents (functional words) should be similar to most languages.<u></u><u></u></span></p>


<p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">Best<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">George Mikros<u></u><u></u></span></p><p class="MsoNormal">


<span style="font-size:11.0pt;color:#1F497D"><u></u> <u></u></span></p>

<p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">[1] Mikros, G., Hatzigeorgiu, N., & Carayannis, G. (2005). Basic quantitative characteristics of the Modern Greek Language using the Hellenic National Corpus. Journal of Quantitative Linguistics, 12(2-3), 167-184. doi: 10.1080/09296170500172478<u></u><u></u></span></p>


<p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D"><u></u> <u></u></span></p><p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">____________________________<u></u><u></u></span></p><p class="MsoNormal">


<span style="font-size:11.0pt;color:#1F497D">George K. Mikros<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">Associate Professor of Computational and Quantitative Linguistics<u></u><u></u></span></p>


<p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">Department of Italian Language and Literature<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">School of Philosophy<u></u><u></u></span></p>


<p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">National and Kapodistrian University of Athens<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">Panepistimioupoli Zografou, GR-15784<u></u><u></u></span></p>


<p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">Athens, Greece<u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">Tel: <a href="tel:%2B30%20210%207277491" value="+302107277491" target="_blank">+30 210 7277491</a>, <a href="tel:%2B30%206976111742" value="+306976111742" target="_blank">+30 6976111742</a><u></u><u></u></span></p>


<p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">Email: <a href="mailto:gmikros@isll.uoa.gr" target="_blank">gmikros@isll.uoa.gr</a>    <u></u><u></u></span></p><p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D">Web: <a href="http://users.uoa.gr/~gmikros/" target="_blank">http://users.uoa.gr/~gmikros/</a>   <u></u><u></u></span></p>


<p class="MsoNormal"><span style="font-size:11.0pt;color:#1F497D"><u></u> <u></u></span></p><p class="MsoNormal"><b><span style="font-size:10.0pt">From:</span></b><span style="font-size:10.0pt"> <a href="mailto:corpora-bounces@uib.no" target="_blank">corpora-bounces@uib.no</a> [mailto:<a href="mailto:corpora-bounces@uib.no" target="_blank">corpora-bounces@uib.no</a>] <b>On Behalf Of </b>Alexander Osherenko<br>


<b>Sent:</b> Monday, October 10, 2011 2:23 PM<br><b>To:</b> <a href="mailto:corpora@uib.no" target="_blank">corpora@uib.no</a><br><b>Subject:</b> [Corpora-List] Are frequency lists of the most languages equivalent?<u></u><u></u></span></p>


<div><div></div><div><p class="MsoNormal"><u></u> <u></u></p><p class="MsoNormal">Hi all,<u></u><u></u></p><div><h3 style="margin:0cm;margin-bottom:.0001pt;vertical-align:baseline;border-style:initial;border-color:initial;outline-width:0px;outline-style:initial;outline-color:initial;font-style:inherit">


<span style="font-size:12.0pt;font-family:"inherit","serif";color:#333333;background:white"><u></u> <u></u></span></h3><p style="margin:0cm;margin-bottom:.0001pt;vertical-align:baseline;border-style:initial;border-color:initial;outline-width:0px;outline-style:initial;outline-color:initial;font-style:inherit">


<span style="font-size:10.0pt;font-family:"inherit","serif";background:white">I am wondering if frequency lists of the most languages can be considered as equivalent. For instance, consider an English frequency list such as the BNC frequency list (<a href="http://www.linkedin.com/redirect?url=http%3A%2F%2Fwww%2Ekilgarriff%2Eco%2Euk%2Fbnc-readme%2Ehtml&urlhash=KPiq&_t=tracking_anet" target="_blank"><span style="color:#006699;border:none windowtext 1.0pt;padding:0cm;text-decoration:none">http://www.kilgarriff.co.uk/bnc-readme.html</span></a>) and a German frequency list (<a href="http://www.linkedin.com/redirect?url=http%3A%2F%2Fgerman%2Eabout%2Ecom%2Flibrary%2Fblwfreq01%2Ehtm&urlhash=99CW&_t=tracking_anet" target="_blank"><span style="color:#006699;border:none windowtext 1.0pt;padding:0cm;text-decoration:none">http://german.about.com/library/blwfreq01.htm</span></a>). The English frequency list starts with the definite article "the". The German one - with the definite article "der". Hence, the literal translation of the word "the" in German will result the word "der".<br>


<br>Of course, it is not always enough to translate directly. However, I wouldn't wonder if say 70%-80% of the most frequent words in the most languages can be considered as equal. Notice I don't say the words should be also ordered in the same manner. For example, word "of" always comes before the word "appear". Nevertheless, I anticipate that words "of" and "appear" are present in the most frequent words of the most languages in every possible order even if particular language uses the word "appear" more often than the word "of".<u></u><u></u></span></p>


<p style="margin:0cm;margin-bottom:.0001pt;vertical-align:baseline;border-style:initial;border-color:initial;outline-width:0px;outline-style:initial;outline-color:initial;font-style:inherit"><span style="font-size:10.0pt;font-family:"inherit","serif";background:white"><u></u> <u></u></span></p>


<p style="margin:0cm;margin-bottom:.0001pt;vertical-align:baseline;border-style:initial;border-color:initial;outline-width:0px;outline-style:initial;outline-color:initial;font-style:inherit"><span style="font-size:10.0pt;font-family:"inherit","serif";background:white">Alexander<u></u><u></u></span></p>


</div></div></div></div></div></blockquote></div><br>

</div>

</div>

</div></div>