[Lingtyp] IPA for linguists' names

Mattis List mattis.list at lingulist.de
Mon Sep 7 07:46:07 UTC 2026


Two points here. First, if you want to archive this project, Yuri, I 
suggest to upload the data after collection to a public repository. I 
suggest Zenodo.org, and share the DOI here, but please also add a clear 
license that would guarantee the re-use of the data. Licensing is 
extremely important to make sure people could potentially remix this in 
future projects that might take up the idea by Martin and provide IPA 
for languages as well.

Second, using Google Sheets is also excluding scholars from China. An 
open server solution would be better. I mention this, since it has not 
been brought up so far. But if the idea is anyway to collect data and 
then make a dump, this should suffice for now.

Anyway: the archiving with a repository seems important, to guarantee 
other people may build on it.

Best,

Mattis

On 06.09.26 21:18, Yuri Koryakov via Lingtyp wrote:
> Dear Jocelyn,
> 
> On 06.09.2026 13:39, Jocelyn Aznar via Lingtyp wrote:
>> Dear Yuri,
>>
>> Le 06/09/2026 à 07:25, Yuri Koryakov via Lingtyp a écrit :
>>> On the one hand, all information not added by people themselves is 
>>> taken from public online sources; people add further information 
>>> voluntarily. On the other hand, this table does not contain any 
>>> sensitive personal information – no emails, place of residence, place 
>>> of work, position, etc.
>>
>> Note that you suggest to people to provide email address in the 
>> comment section.
> 
> Yes, you're right. Although this only applies to quite rare cases where 
> there are complete namesakes and they need to be distinguished somehow, 
> and only as one of the methods. I've added ORCID, by the way, thanks for 
> the idea!
> 
>> In some countries (like the US), the pronoun to use is actually a 
>> sensitive personal information, as it reveal gender info. Native 
>> language can be a proxy for ethnic identity which also can be 
>> problematic in many countries. Google Doc, especially with a free 
>> account, is not safe for personal info, not to mention that someone/AI 
>> could vandalize the spreadsheet and insert/modify the data, and when 
>> collecting personal info, the database creator is responsible of 
>> providing a database which is secured enough according to the 
>> sensitivity of the info collected.
> 
> Yes, it's very important. For a while, I'll keep an eye out for 
> vandalism by checking the edit history.  After the active phase of 
> adding names, I will make a static webpage and change the way new names 
> are added.
> 
>> More practical issue, I can imagine in the US that inviting a 
>> researcher that asked for the pronoun they could be problematic, 
>> either for the researcher or the lab inviting them. As it is 
>> information available on Google Doc, this will be automatically 
>> scanned by the IA that the border police uses.
> 
> Horrible, I didn't know about such problems in the US (I would 
> understand if it were Russia...).
> 
>> Birthday can also be problematic actually, as the other 
>> disambiguators, as they allow to identify the person, which is normal 
>> for the purpose of this database, two people having the same name 
>> might want different ways of being addressed. Better in this case 
>> ORCID or professional URL which are publicly available data.
>>
>> Somehow, GDPR might seem annoying, but, in this day and age of IA, any 
>> info available on internet is being scrapped and compiled, and 
>> automatically analyzed by more and more fascist governments and 
>> organizations for whatever reasons they have, or will have. Somehow, I 
>> wonder more and more about the relevance of open access data in this 
>> new context.
> 
> To sum up, I can say that all sensitive information is added by people 
> themselves. Just in case, I've added the following warning: *Warning*: 
> All information you provide will be publicly available. If you feel any 
> information is too sensitive and/or could be used against you, do not 
> provide it.
> 
>>>> - There isn't only one valid pronunciation of my name.
>>>
>>> The IPA column allows you to specify multiple pronunciation variants, 
>>> both equivalent and language-specific.
>>
>> Right.
>>
>>>>
>>>> - I'm ok with people having difficulties pronouncing my name (at the 
>>>> end of the day, no one can master the pronunciation of the many 
>>>> phonologies we work with in our profession).
>>>> - Pushing people to pronounce your name following your expectations 
>>>> can cause situations in which they feel really ill at ease because 
>>>> they cannot pronounce your name.
>>> Well, we don't force people to pronounce the names exactly this way, 
>>> but we do give them a reference point. For example, when I just see 
>>> the name Jocelyn, I don't even know what the first sound might be – 
>>> [ʤ / ʒ / x / j] or something else.
>>
>> Of course I understand this issue.
>>
>>>> - I'm also ok with people adapting my name if it is too hard to 
>>>> pronounce, that's usually what happens in practice.
>>>> - From a more technical angle,it would make more sense if the IPA 
>>>> transcription follows the same order of the first name + other name.
>>> Okay, maybe. But this way one can show the order in which a person 
>>> prefers to pronounce their name. (By the way, the most common order 
>>> in this column is: first name + other name.)
>>
>> This is not so clear from the instruction neither, and this is not 
>> what's systematically done from the data that was already populated.
> The Guide says "Please provide a transcription of your full name (*in 
> your preferred order*)."
>>
>> For the purpose of this database, I would have a dedicated column for 
>> the preferred name for a seminar/event, not just "Other names" which 
>> can lead to confusion as it might contain many info in different 
>> order. 1 info : 1 column.
> If I created a separate column for each type of information, the 
> database would grow too large. And in general, a person can provide the 
> preferred name for a seminar/event when registering for that seminar/event.
>> I think this project would be better handled through the server of an 
>> institution (well, more of our institutions rely on Microsoft/Google 
>> and co., but if they pay for the services, the data might be actually 
>> properly handled), with a better interface, like a form in which 
>> people can enter these info and a page that displays those info.
> Yes, perhaps that would be the ideal solution. But I don't have such 
> possibility (and I'm not sure all linguists would be willing to provide 
> such information on a Russian institution's server).  I proceed from a 
> simple maxim: The best is the enemy of the good.
>>
>> Best,
>> Jocelyn
> Yours,
> Yuri
>> _______________________________________________
>> Lingtyp mailing list
>> Lingtyp at listserv.linguistlist.org
>> https://listserv.linguistlist.org/cgi-bin/mailman/listinfo/lingtyp
> 
> _______________________________________________
> Lingtyp mailing list
> Lingtyp at listserv.linguistlist.org
> https://listserv.linguistlist.org/cgi-bin/mailman/listinfo/lingtyp



More information about the Lingtyp mailing list