Estrella's Lab Notebook

Estrella's Lab Notebook

SUMMER 2026 SCHEDULE (WIN Internship):

MWTHF: 8-5pm

34 hours

T: 2-5pm

3 hours WFH

Direct supervisor

Stephanie Grasso Diana Cruz Camille Wagner Rodriguez

Table of contents

We don't have a way to export this macro.
We don't have a way to export this macro.
We don't have a way to export this macro.
We don't have a way to export this macro.
Button | Mosaic

Longitudinal PPA “Picnic” Project Reliability

Date Priorities Notes @Camila Maldonado to practice running whisper on these samples: by @Camila Maldonado to note down questions/errors encountered (if any) to discuss with @Estrella Palomo on @Daniela Ortiz-Aviles to look over Whisper wiki page to familiarize herself with the content: -Worked on finalizing trimmed acoustic audios for Spanish → Errors received. Asked in supervisor chat pending next steps. -Trained Camila on Whisper -Added whisper practice to Camila’s notebook on Box (Also made this an action item for her) -Checked training quiz with Arely to ensure it was ready to administer -Reached out to Daniela to set a time for whisper training next week.
Date Priorities Notes -Ran the Catalan files through CLAN + uploaded outputs to Box under corresponding folder Encountered a few difficulties with assigning the directories but was able to complete the task -There were no nodes available so I was not able to run the pipeline, but I will try again next week. -Helped Jada get oriented with the steps to running the linguistic pipeline -Continued DEMQoL transcription CGSG013 Question -@SG, given that Jada and I overlap a great majority on Monday (from 11-1:30) when would be the best time for our weekly meetings? Thank you! -Clipped BISD025 in Catalan -Clipped BISD026 in Catalan (PENDING) -Talked with Arely about training
Date Priorities Notes -For Helena: Let Helena know to prioritize Important Event past the WAB and control samples → Filter smartsheet for both languages and timepoints Pre, 12m, Post, and then the following timepoints @Estrella Palomo to filter CS smartsheets for Helena to reflect priority (Important Event) Hierarchy: Pre, 12m, Post, and then remaining timepoints. -Send message to Helena -Finish drafting training plan -Sent Helena a message on what transcriptions to prioritize -Finished drafting training plan for students (in my lab notebook) -Clipped BISD025 -Clipped BIOBS023 -Clipped BISE013 -Clipped BISE022 -All changes on Box and Smartsheet -Worked on poster presentation Before meeting for training, student will read through the Wiki pages and come prepared with questions. Student team for clipping/whisper Fall 2026: @Arely Aguilar, Gabriella, @Sadie Yarbrough, Camila (pending confirmation from Diana) @Estrella Palomo to send message on Student Team channel when all individuals have been added to Teams by and to let @Arely Aguilar know that you will send that initial message Training: Based on what you read from the wiki pages explain to me: -How to identify what needs to be clipped  -Where clipped audios go -How to identify what audios need to be processed through Whisper -Where to allocate audios that have been processed through Whisper  -Where Whisper outputs should be housed  -Where you indicate the status of clipping/whisper -What does it mean for a patient to have been administered Set 1 + 3/Set 2 + 3? Review folder structure of Connected Speech Data: Clinical Trial, Observational, Screening + Basic Eval, and Controls + How to identify what category the sample belongs to Review how to identify if an audio is Pre-R01 or R01 using the smartsheet. Once this has been done → Show student where to find the Pre-R01 samples and how to properly organize them on Box Explain the importance of proper file naming and proper storage. Emphasize that cache must be cleared after every whisper run. Assign student whisper/clipping practice  Any questions? Short Quiz: “Using the smartsheet as a guide, can you show me how you would find this patient’s audio on Box?” Do it for R01 and Pre-R01 and also ask “Where would you put the audio once it has been processed through whisper and where will the whisper output go?” @Estrella Palomo to incorporate questions regarding Whisper into the existing CS quiz (or make a separate one) Volunteer + Makeup Hours: Mon August 3rd- 8:00-10:00 (2) Tue August 4th- 8:00-2:00 (6) Wed August 5th- 8:00-2:00 (6) Mon August 10th- 8:00-10:00 (2) Tue August 11th- 8:00-2:00 (6) Wed August 12th- 8:00-2:00 (6) Aug 14-Aug 22= OOO traveling -Organized project manuscript folder @Estrella Palomo to ensure transcripts are copied over (not moved) into the 1. CLAN Transcriptions_YYYYMMDD AND 2.Stripped transcripts folders by @Estrella Palomo to add the final output Kesha shared along with file Estrella generated for Linguistic pipeline into 2. Stripped transcripts/Output folder by Next steps for project: @Estrella Palomo move these next steps to the Long Project Page in a place that makes the most sense by Resolve missing MMSE scores Resolve that some participants have more than two visits and we need to select only two visits per participant @Estrella Palomo to run non-stripped samples through CLAN to extract those variables by @Estrella Palomo to run aby coustic pipeline on the audio samples by Stephanie and Estrella to review reliability values for her cohort so we can have that clearly outlined Stephanie and Estrella to review other pending action items in this file for analyses later Action_Items_Full_Cohort_Inclusion Kesha steps: reordering variables so they match across Catalan and Spanish for Linguistic pipeline Kesha also checking that variables missing in Span or Cat are meant to be Future analyses: Run analyses within each variant separately to see outcomes and speak with Dr. Santos and Núria to confirm biomarker status, potentially restrict analyses within each variant to those with biomarker confirmation. Run analyses with and without MMSE to also understand influence of cog factor as covariate, compile bilingualism questionnaire to explore those variables, include neuropsych measures for more comprehensive Table 1. -Began to draft plan for training students on Whisper -Helena is caught up across all timepoints → What is next priority? -Finished working on WIN poster -Recorded cache clearing tutorial with Kesha -Added in reason for sample exclusion under “Notes” column in the 0. Master Data Sheets. -Worked on organizing project folder, need clarification on the following: Within the transcription folder, what is supposed to be housed in the “Transcriptions” vs “Output”? Are the Whisper transcriptions included in either folder? Or is it the linguistic pipeline outputs? Within audios, I have the same question: what are the “Outputs” meant to be? Is this where Whisper outputs are meant to be located? What is the date YYMMDD for each folder going to be? -Check-in on the samples BILP010 and BILP020 Obs1 in Catalan for date discrepancies as pointed out by Claude. Fixed the dates on the parent Catalan sheet. Month and date had been switched and were written as “2023-05-12” instead of “2023-12-05”. Notes from meeting with Dr.Grasso: Fillers: svPPA had a decline over time in “fillers” lvPPA had a very slight increase and nfvPPA also a slight increase but not as much as lvPPA Offsets -Family_Variable_Definitions_Table, print + review -AoA: Higher for semantic for some reason! -Was offset included in samples from yesterday = No, now it is accounted for and includes visit 1 to visit 2 instead of days apart which hid some of the reuslts -Reformat significnat/not significant figure -On top of the graphs include which family of measures it belongs to and include a bullet point summary under each one of the main findings. in conclusion/discussion discuss findings and how this relates or deviates from other literature on connected speech findings. don’t forget to add in the graph that gives more insight on the participants (See comment on the poster) Mess around with a title + rework the co-authors + affiliations to fit into poster. -Enumerate action items by topic and take a look at everything on Monday with Dr.Grasso -Also, check-in with Dr.Grasso about what work can be given to Helena since she has caught up with control clipping
Date Priorities Notes @Estrella Palomo To work on Diana’s CGSG013 transcription for DemQoL -Clipped BIOBS009_Cat -Clipped BIOBS011_Cat -Clipped BIOBS022_Span Pt.2 -Clipped BIOBS022_Span Pt.2 -Clipped BIOBS014_Cat Pt.1 -Clipped BIOBS014_Cat Pt.2 -1/2 Done clipping BIOBS013 -Smartsheet and Box updated with audios! -I have oriented Helena on where she can find the audios that need to be transcribed by providing her a link to the smartsheets. I also added the “Dx” column to indicate “control sano” for further clarification. -Tomorrow I will attempt to run the linguistic pipeline locally (if TACC is back up tomorrow I will run as normal). @AG would you be able to quickly orient me on logging in to the computer with the local Kesha scripts? Or if there is a wiki page on running the pipeline locally could you send it my way? I appreciate it! -@SG, I will continue working on the introduction + methods tomorrow and will send a rough poster draft tomorrow. Lit. Review Measures Spreadsheet: Link to google doc: (I am drafting what will go on the poster here) Link to poster draft: Note: I have to choose a style from here, waiting to have a better idea of how much space will be needed. @Estrella Palomo to review participant videos and note overlaps in symptom manifestation across the 3 variants + map out similarities/differences between what is seen in videos and what has been presented in literature review by -Draft a message to Helena to remind her about the due date of July 14th for the transcriptions in order to begin running the linguistic pipleline. Send on Friday. -Clipped BIOBS018 Pt.2 -Clipped BIOBS019 Pt.1 and -Clipped BIOBS019 Pt.2 -Clipped BIOBS021 -Clipped BIOBS022 Pt.1 -Clipped BIOBS022 Pt.2 All changes are reflected on Smartsheet and Box. Notes on schedule changes: Next Tuesday WIN required events will be wrapping up sooner. Thus, I will be making up my hours from this week on Tuesday, July 14th. I have made this change on the calendar. Thank you! -Final verification of dates on REDCAP. Pending dates for PTSG001 AND PSG002 dyads due to file naming. Will clarify in meeting. -Continued reading through paper -Run Whisper for priority samples → Finished running all Spanish samples that were pending. Uploaded to Box and smartsheet has been updated! -Clipped BISD013_Obs2 -Clipped BIOBS017_Obs1 -Clipped BIOBS018_Obs1 -1/2 done clipping BIOBS019_Obs1 Discussed with Dr.Grasso: -Controls will now be prioritized for clipping + transcription in order to begin completing this subgroup. The controls can be found on each parent smartsheet using the Dx column and labeled as “Control sano”. Note to self: Send a message to Helena to inform her of this new priority. Let her know that these transcriptions will be the next priority after the WAB samples but do not have a set deadline! -Worked on adding correct date information in registro instrument. Hello @AG would you be able to meet tomorrow or Wednesday for a few minutes regarding a question about a date for a patient? -Finished VISTA training → Changes on excel located in Box training folder -Clipped BILP020_Obs2 in Spanish –> Updated Whendy on status + also provided Catalan sample -Located BILP020_Obs1 which = BILP020_12m, audio had been clipped. I made sure to copy over the Obs1 audios to the clinical trial folder and indicated “12m” timepoint. Added comment to smartsheet. -@SG since BILP021_12m has no audio, I have let Helena know to not worry about transcribing this sample + to not transcribe BILP020_12m as a transcription already exists. -Note on BILP021: Filled out rows to indicate “Administered but cannot be located”. Smartsheet comment already exists to indicate that there is not audio. -Continued reading through papers. I finished “Written picture descriptions distinguish variants of primary progressive aphasia” today and began reading “Reported symptoms and patterns of language impairment in bilingual speakers with primary progressive aphasia: a retrospective study”. -Note on BILP020 (and similar scenarios): Problem: There was no BILP020_12m in the Spanish Parent Sheet but all other timepoints existed (3m,6m, pre, etc.) The patient also did Obs1 and Obs2. In this case, the Obs1 is treated as the 12m AND as Obs1! So when clipping, label as 12m for the Clinical Trial folder and as Obs1 for the Obs folder. -Clipped BISD018_12m Catalan -Clipped BISE010_12m Catalan -Clipped BISD019_12m Catalan All audios have been uploaded to Box and the smartsheet has been updated for clipping reports + Helena’s temp report. -I started to format the spreadsheet we discussed in our meeting yesterday for the measures. I will be adding them in and hoping to have everything by Wednesday. -I made good progress on the VISTA training, I expect to finish it no later than Monday as I am more than halfway done. -As for the audios, I am nearly done clipping I am missing 2 audios. However, BILP021_12m has a clinician note that the audio cuts off mid session. After reviewing the videos, none of the WAB samples have audio. What should I do about this? Other than that I need to clip BILP020_12m in Spanish but I can’t find it on the parent sheet or on Box. -Finish clipping SSLP007_12m Span -Start VISTA training assignment -Continue clipping priority samples -Working on VISTA training -Clipped SSLP007_12m Span -Clipped SSLP007_Pre Span -Continued reviewing patient videos for the different variants -Clipped BIOBS003_Obs2 Catalan -1/2 done with BISD019_12m in Catalan @Estrella Palomo to run linguistic pipeline by and concurrently draft intro + methods of poster. @Estrella Palomo and @Stephanie Grasso to meet the week of July 13th @Estrella Palomo and @Stephanie Grasso to meet and run analysis for WIN poster on -july 13th → meet with Dr.Grasso -july 14th → Helena completes transcripts for priority audios july 15th-17th → run pipeline and work on intro and methods july 20th → poster july 23rd and 24th → analysis with Dr.Grasso -Met with Katie to discuss Qual Analysis -Clipped BISD010_12m Span -Clipped BISE010_12m Span -Clipped SSLP006_12m Span -1/2 Done with SSLP007_12m Span -Watched patient videos + took notes on how the different variants displayed symptoms -Took notes on how measures are defined in Canu paper -Took Bilingualism Survey -Clipped BILP020_Obs2 in Catalan -Clipped BILP021_12m Part 1 + 2 in Spanish -1/2 clipping BISD010_12m in Spanish -Sent Helena message on where to find priority samples @Estrella Palomo to meet with Dr.Grasso to discuss file localization @ 3pm @Estrella Palomo to prioritize clipping the Catalan and Spanish samples that were missing from June 26th search. @Estrella Palomo to copy over transcriptions of Cat and Span samples over to project folder. @Estrella Palomo to add 2 columns to each sheet (Cat and Span) indicating if the sample was clipped again + if aligment with transcription was verified. @Estrella Palomo to confirm meeting date and time with Katie (currently set for ) -Added the following columns to spreadsheet: “Audio Copied?” → Audio file was copied over to my project folder on Box “Needs to be Clipped?” → Indicated if audio is pending clipping + is an indicator to look at column S “Transcription Copied?” → Transcription file was copied over to my project folder on Box Verified with Transcription?” → Added but as all transcriptions were accounted for it is not needed. -Went through each participant timepoint again to ensure all audios were accounted for in my project folder. The ones that require clipping are indicated “Yes” on the “Needs to be Clipped” Column. -I copied over all of the transcriptions for the audios in Spanish and Catalan. All were accounted for except for audios that are pending to be clipped. -Clipped Spanish sample for BIOBS003_Obs2 Notes: Participant spoke primarily in Catalan. There was a lot of clinician speech to be clipped out. -1/2 done with BILP020_Obs2 in Catalan. -Finished reading Hardy paper -Checked that variant across the timepoints for BISD007, BISD012, and BISD013 are classified as svPPA. Note: After finishing BILP020, I will be focusing on getting the Spanish audios done first as Catalan audios take me longer to get through. Updated Clipping Priority: SPANISH: BIOBS003_Obs2 Done BILP021_12m Done BISE010_12m Done BISD010_12m Done BILP020_12m → Need to find SSLP006_12m Done SSLP007_Pre + 12m Both Done CATALAN: BIOBS003_Obs2 Done BILP020_Obs2 Done BISD019_12m Done BILP021_12m → No audio BISD018_12m Done BISE010_12m → Done @Estrella Palomo to copy over Catalan and Spanish Picnic samples to project folder. @Estrella Palomo to look through Qualitative Analysis Training Page + login to Atlas.ti Tasks done today: -Copied over Span and Cat samples to project folder and verified if all Picnic Clips were accounted for. I will be prioritizing the samples below to clip + will do so alongside the transcriptions to ensure the new audio clipping aligns. SPANISH: BIOBS003_Obs2 BILP021_12m BISE010_12m BISD010_12m BILP020_12m CATALAN: BIOBS003_Obs2 BILP020_Obs2 BISD019_12m BILP021_12m BISD018_12m BISE010_12m -1/2 done reading through Hardy paper on Symptom-led staging -Began to watch videos of the different variants, started with logopenic. -Checked out lab laptop with Diana to work remotely. -Read through Qual Analysis Training and took note of a few questions for my meeting with Katie -Note on Catalan 12m samples: BISD018 → located and copied to Catalan clinical folder BISE010 + BILP021 (Pre-R01 Cases) → Go to video source and clip. -Note on Catalan Obs2 sample: BISE004_Obs2 → Not administered! The following audios are to be prioritized and clipped: SPANISH: BIOBS003_Obs2 BILP021_12m BISE010_12m BISD010_12m BILP020_12m CATALAN: BIOBS003_Obs2 BILP020_Obs2 BISD019_12m BILP021_12m BISD018_12m BISE010_12m NOTES: IMPORTANT!!! Verify clipping with the transcriptions and ensure start and end align with what is reflected on the transcript. Note: Check for any missing transcriptions while copying over, also add this column to designated sheets. On the sheet, also indicate with a column if the audio has had to be clipped again. This is to ensure that the audio was checked and aligned with the transcription. “100- ALL Audios IP Sessions” → In the event that there is any uncertainty about the administration of a timepoint check within this folder. It may be that a clinician did not copy the audio over. @Estrella Palomo to search for patient samples BILP008, BILP030, and BIML003 in Catalan by @Estrella Palomo to search for patient sample BILP020 in Spanish by -Finished an additional paper this morning (Costa) and finished up reading Mueller paper I had started. -Compiled notes on all the papers I read and summarized everything in a table in both my lab notebook and the Longitudinal PPA Wiki -Attended Support Group Meeting -Met with Dr.Grasso -Met with Ana Pau to orient myself with the assigned task and navigating REDCAP -Searched for picnic scene samples that SHINY App indicated missing -Send Sadie picture + description for post! Tasks done: -Reclipped audios -Read through Hardy paper -Clipped BISD022 -Clipped SSAD004 -Clipped BISE027 -Clipped BISE015_Pre -Clipped BISE015_Post -Wrote blurb + sent headshot to Sadie @SG I just wanted to confirm my attendance for the support group meeting tomorrow as we had discussed last week. I see that there is a zoom link on the calendar notes, is that how I should join? Thank you! Question about BISE015: -In the Pre session MAINCat + Sunday tasks were administered. In Post, MAINDog was administered. Both labeled under Set 2 + 3. Is it normal for same patient to be administered different sets? The primary clinician is listed as Jan Holst. Estrella to look at Redcap instrument for connected speech to see if things look congruent. We looked in our meeting and added a note about this. @Estrella Palomo to take Bilingualism survey to see the structure https://redcap.dellmed.utexas.edu/surveys/?s=NaVHkvSJbsdHoGga and to read the following website https://blp.coerll.utexas.edu/ by @Estrella Palomo to look together with @Stephanie Grasso at which samples need to be prioritized for transcription given that Whendy is slowing down. We can jointly message Jaume as well by -Finished clipping BISD012 -Finished reading + annotating Canu paper -Started to read Hardy paper on lvPPA -Finished Clipping BILP030 Part 1 + 2 -Finish CGSG009 transcription -Clip audio? -Continue reading through papers → Add table to the project page -Finished CGSG007 Transcription -BISE029 Communication Partner in English -Clipped BISE014_Obs2 -Clipped BISE014_12m -1/2 done clipping BISD012 -Working through papers on lab Zotero currently reading through Canu paper Notes on Patient: -Adhered to guide learned during therapy -Could be understood, although there was some delay between words -Ask Dr.Grasso about working similar hours to last week for a half day next Friday (July 3rd) because parents coming to visit. Work until 6 T and TH! -Finish CGSG007 transcription -Start CGSG009 transcription -Clip audio -Finished CGSG009 transcription -1/2 of CGSG007 transcription finished -Clipped BISD011 audio -Finalize the transfer of PTSG012 Transcription to REDCAP -Clip BISD021_6m - Tasks Done: -Finished CGSG010 transcription → Took incredibly long! Ended up being 13 pages as the interview was almost 1 hour long! -Clipped + Ran whisper for CGSG009 and CGSG007 Semi Structured Interview -CGSG009 Transcription → IN PROGRESS. I am about halfway done. Will finish tomorrow. -Clipped BISD021_6m → Box and Smartsheet updated -Searched + found CS administration dates for samples Dr.Grasso sent. Reminder: I will be WFH tomorrow from 9-1! Located BISE005. BISE006, BISE010, and BISE011 all in Pre-R01 videos processed through whisper. None of the audio files or tasks had any names on them. I resorted to looking through participant videos to determine the date of administration for tasks. BISE004: July 15, 2024. No audio file existed within the connected speech folders. Located Obs2 video under R01 participant videos here: https://utexas.box.com/shared/static/9zegfgj9vv033izy1k25i40gxiucn79l.mp4 BISE005 → Picnic Scene 03/15/21 @ 33:40 BISE006 → Picnic Scene 07/05/21 @ 29:25 BISE010 → Picnic Scene 01/26/22 @ 48:00 BISE011 → Picnic Scene 04/06/2022 @54:00 -Review Katie’s response to questions + Start on a new transcription -Finish clipping SSAD001 -Run Spanish tasks through Whisper -Add table to project page -Finished PTSG010 Transcription -Finished PTSG012 Transcription -Finished CGSG008 Transcription -Katie clarified where to store audio and whisper output within Box. Also mentioned the addition of new protocols: Whisper will run for the interview portion only. This means to clip the appropriate session to save time. There is now an unintelligible code in the protocol. -Ran Whisper for pending Spanish samples. Uploaded to Box and Smartsheet updated. Clipped ESAD001 →FINAL CLIPPED Audacity file in both Spanish and English folders. All tasks uploaded to Box + Smartsheet updated. Questions: Ask Katie about the other person who was speaking in PTSG012 → What is the name replaced with? Unsure if this is the care partner or not. -Ask Dr.Grasso about WFH on Tuesdays due to commute time. Tasks Done: Clipped BIML004_Post. Uploaded to Box + Smartsheet updated. Began to clip ESAD001_Post → Confirm audacity file location Edited “Troubleshooting” tab in Student Lead Wiki -Clipped the following audios (all in Spanish): SSAD004_PRE BILP038_Pre_Part 1 BILP038_Pre_Part2 BISD026_Pre BISD022_Post BILP037_Post Started to clip: BIML004_Post ESAD001_Post, can’t find Post Tx audio containing WAB. This is the one that was pending to be clipped here: There is a comment indicating only Picnic and Important Event were administered. Primary clinician Diana. Attempted to clip BISE027_Span_Pre but there were no WAB samples in the audio. Only English under “participant videos” do I indicate “Not administered” in the Spanish smartsheet? Ana P primary clinician -Check for any messages from Whendy/Helena regarding file localization -Finish last transcript (check if Katie answered questions/comments sent in QUAL Analysis Chat) -Ask Camille about any outstanding questions regarding the VISTA training. -Continue reading through RESPALDO. -Finished all transcript assignments for Qual Analysis. Today I worked on PTSG013. I ran the video through whisper as REDCAP was empty. Pending new assignments and confirmation on a few questions. All links accessible on Smartsheet. -Continued VISTA training -Nearly done reading through RESPALDO paper -Clipped BISE028_Pre → Uploaded to Box + Smartsheet updated. -Began to clip SSAD004_Pre → Will finish on Monday -I was going over the clipping reports and noticed that a large chunk of Pre-R01 12m and 6m videos were marked as “Could not be located”. I began to sort through the files and linked the WAB samples in the comments for the ones I was able to find. I will also finalize this by Monday to make future clipping easier. -On that same note, I had added a “Troubleshooting” tab on the lead wiki page. I will be updating it now that I have gained a bit more proficiency in locating missing samples. -Run Whisper for SSLP006, SSLP010, and SSOBS008. Send Whendy Box links when completed. -Continue next transcription for Katie -Begin literature review for RESPALDO, Manuscript, and Literature Review -Ask about adding a section that emphasizes the importance of accurately updating the smartsheet. Also, include a note about sessions that were divided into 2 sessions to ensure none are missed. This would go in the clipping process wiki. Ran Whisper for PRIORITY samples: SSLP006, SSLP010, and SSOBS008 But also, for Spanish clinical and Pre-R01. Finished transcription for PTSG010. Ran video through Whisper because REDCAP was empty. Linked here: Met with Camille for Reliability training -Links provided: VISTA Summary: VISTA Reliability R2 (training materiales included in this wiki): Sign in to MADR LAb toolkit: https://comm-sg-p01.moody.utexas.edu/accounts/login/?next=/ -Watch the POM video training by ! Meeting notes with CWR: Phase 1 or 2? No!!! For monolinguals there are no phases. Untrained vs trained… trained means it was worked on during the treatment In materials that CWR will send, there are groups of links that will be helpful. In regard to POM training video, second timestamp is a bit more complicated but it’s the actual procedure process. Pay extra attention! -Message sent to Whendy regarding Whisper: Hello Whendy N Avila Motta,  Here are the links to the Whisper outputs SSLP006: SSLP010: SSOBS008: -Message sent to Katie with questions/comments about Transcription procedure: Hello Cole, Katie!    Here is a summary of the questions/comments I have:   When running the video through Whisper are we to run the entire audio or clip to the appropriate spot and then run through Whisper? For PTSG010, I clipped based on the timestamps that were provided in the comments and then ran through Whisper. I allocated the clipped version of the audio here: . I did this to decrease the time it took for the video to be processed since the interview itself was only ~7 minutes. I have put the whisper output within the same folder as the audio here: I may just be overlooking it, but I don't see a whisper column in the SSG_Transcriptions smartsheet that are indicated on the Wiki page 1. Who Processed Via Whisper? 2. Video/Audio Processed Through Whisper” I noticed while uploading the clipped audio and whisper output that Ana had already uploaded a whisper output. However, the REDCAP was empty. Was there somewhere she indicated running whisper that I missed? One last thing, I agree with Jimena that we'd benefit from having an <unintelligible> code  -Run Whisper for pending samples -Continue transcription of CGSG010 @Estrella Palomo to read RESPALDO manuscript by and come with a few notes indicating what you think the interview data is meant to provide -Clip an audio -Meet with Dr.Grasso @ 11:15 @Estrella Palomo Read the following manuscript to come prepared in next meeting to describe the participant: BISE028 by -Install Rstudio server on the dell computer On Helena’s Temp Report check if BILP007 truly is missing their Picnic Scene description -Establish strong foundation for rationale for the project Literature review for Introduction @Estrella Palomo to add a section to the project page for the literature review, pulling in previously read papers and their findings and adding this in a table format that allows us to digest the most pertinent information into the Longitudinal Project page , add papers to lab Zotero connected speech, spontaneous speech, picture description, primary progressive aphasia, semantic dementia, +/- longitudinal, trajector*, progress*, pattern, decline Hours for next week M,Thu, 8-6 (half hour lunch) W: 8-5 Tuesday 2-5:30 Friday: 9-1 WFH WIN Presentation Completed transcription of CGSG010 BILP007 Picnic Scene PRE-treatment in Catalan → seems to not have been administered. I looked through the videos here: and did not find the picnic scene in Catalan. If I could get a clinician to verify just to be extra sure I’d appreciate it before sending out a confirmation message to Helena! Confirm where experimental room is → 2nd door to left of big office space Completed BISE028 VISTA session as communication partner: -Note: If clinician does not provide the zoom link (e.g. Today), look through the Cronograma spreadsheet: and find the zoom link under “TX”. -Session went well, made the conversation sound as natural as possible for the patient. -Have not received any feedback from clinician (Maria Pinart) → reach out to ask her ways to improve? Assisted Dr.Bao in blindly testing the Cinderella task -As of right now, the task is composed of 4 sections and takes a total of 5 minutes. Section 1: Tell the story of cindrella in as much detail as possible. 1 minute to do so. Section 2: Describe 2 images in as much detail as possible. Section 3: (Without being shown the previous 2 images) recall the details in the 2 images Section 4: Two questions -What were the animal names written on image? -Which of the following pictures appeared in one of the 2 images? Feedback provided to Dr.Bao: -Timing for first task felt rushed (he said the priority is detail not necessarily completion) -Sufficient time for sections 2 + 3 Notes on sanity checking: Go to redcap → search patient (e.g. BILP029) CLICK UNIQUE D nd go to connected speech instrument check preimera and seguimiento pRE R01 patients → MADR participant information smartsheet should suffice for checking sanity For the literature reading: Utilize Zotero to upload into the MADR Lab ”Longitudinal PPA” project. Read through introductions of the paper to get a better sense of what they are about Using Gemini (provided with UT), indicate a prompt to create a tabular format of all the literature with information about the participants in the studies, etc. This will serve as a base to create a centralized spot for the literature. Once summarized, tweak to match Message sent to Whendy regarding file localization:   I have a few updates for some of the files I had previously marked as "Not administered". I went back to verify that the "Completed" audios with Picnic marked as "Not administered" by our clipping team was accurate. After going through each individual audio it seems the task had originally been missed (or not indicated on the smartsheet) when being clipped. Here is a summary of my findings:    SSLP006_Pre: Marked as completed with Picnic "Not administered" I was able to find the task and clipped it here:  Here is the clipped version here:   SSLP010_Pre: Has been found and clipped. Please note: there are two parts to this audio. A student marked the audio as "Completed" but accidentally missed the second part. I will be returning to this audio ASAP as I only clipped picnic scene to get this sent out to you. I will indicate this comment on Smartsheet. Here is the clipped version: SSOBS008_EB1: Had also been marked as "completed" but no picnic scene. I was able to find it. Here is the clipped version:   Below you will find the ORIGINAL audio sources for your reference:   -SSLP006: -SSLP010: -SSOBS008:   There were no nodes available to run the above samples through Whisper. I will prioritize this tomorrow.   NOTE: I noticed on the your smartsheet that BILP013 was marked as "No audio clip". It likely got lost in the thread but this was clipped and ran through whisper last month. I have pasted the links below:   Picnic Scene Clipped: Picnic Scene Whisper:   -Read through Communication Partner VISTA page -Clarify 16:45 BcN time is 10:45 am here. -Continue working through CGSG010 transcription → ask Katie for times to meet? Ask how to indicate a stutter? Time stamp 46:00 for CGSG010_PostPhase2 Nearly done transcribing CGSG010 Reclipped audio Reviewed VISTA page 16:45 BCN is 9:45 am in ATX! Camille booked experimental room from 9:30-10:00! Questions: -What does it mean to be a “communication partner” I signed up for two sessions (June 10th, Castellano and June 22nd, English). Review this page: -How do I receive access to RedCap to assist with transcribing? Support Group Redcap: https://redcap.dellmed.utexas.edu/redcap_v17.1.2/index.php?pid=3772&__record_cache_complete=1 -Clarify hours for working past 5pm when needed. Clarified hours at top of lab notebook and discussed on Re-Review Student Team Lead Tasks: -Meet with Dr.Grasso to discuss: SSLP008_Pre Picnic Scene — Needed clipping (confirm it has been clipped by Arely and link you provided seems to have expired so you'll need to hunt for this one) The full session was located, but it's over an hour long and still needs to be clipped down to the Picnic Scene. Can you walk me through what went wrong with your clipping process for this one? 🔗   BIML001_Post Picnic — Was missing source audio (confirm it has been found) This was run through Whisper, but I can't locate the audio file it was run on. Can you let me know where you pulled the original audio from? It's not in either of these folders: 🔗 🔗   SSLP004_Pre Picnic — Reclip pending since 5/17 Needed to be reclipped (check Arely did it). Please come prepared with a plan for how you will monitor this as the clipping lead! 🔗   SSOBS012_EB1 — Reclip pending since 5/17 You'd started on this one, but the sample looked unchanged on my end — I think the edits may not have saved. Please see what went wrong but Arely I believe re-clipped 🔗 Tasks done today: -Answered Whendy’s messages on the Transcription channel. -Made notes on the sample cases Dr.Grasso outlined in message. -Reviewed “Transcription Manual” → Am I to transcribe the audio from scratch or is there something on RedCap I need access to? -Verified whisper smartsheet was up to date with output from May 28 and Today. -Clipped BILP035_6m -Clipped BISD023_6m -Clipped BISD019_12m All changes reflected on Box and Smartsheet. -SSLP008 → Accidentally uploaded full version instead of clipped version. Error on my end. Updated Whendy on this today! -BIML001 → Found here: Originally clipped it as is, noticed whisper output had clinician speech. Reclipped today. Updated Whendy on this today! SSLP004 →Jada clipped this sample and I ran it through Whisper and the output could be found here: . I must have missed the reclip request as it was sent 5 pm 05/27, my last day before returning for WIN was 05/28 where I prioritized uploading whisper output, updating smartsheet, and clipping samples that I had been unable to locate until that day. I also reclipped a sample that day for Whendy, I overlooked the request for this though. SSOBS012 → I reclipped the sample as Whendy asked and it can be found here: . I uploaded this on May 28. It ends at the timepoint she indicated to me.
Date Priorities Notes -Clipped picnic scenes for SSSLP008 and SSSE009. I also ran the samples through whisper. -Uploaded the whisper outputs and updated the smartsheet! -Began to clip BILP035 → Will have to continue this once I am officially back in lab. -Ran whisper for all Catalan samples! -Jada searched for BISD007, Camille confirmed pre picnic was NOT administered. -Met with Dr.Grasso, discussed continuing clipping and running Whisper. -Add specifications to the breakdown of the options for RELY spreadsheet. -SSSLP008 and SSSE009, need to confirm picnic administration. Audios are really long, asked if there was another way to confirm administration (such as clinician notes). -Clipped BILP018 -Ran BILP018 and BILP010 through Whisper -Difficulty locating BISD007, sent question in supervisor chat -Met with Dr.Grasso: -Continue running Whisper for priority samples, TACC should be up and running and can be accessed now that we are done with finals. -Prioritize clipping the samples for Whendy (BILP010, BILP013, and BILP020) and make sure -Started clipping BILP013 @Estrella Palomo Troubleshoot location of files Whendy is searching for. @Estrella Palomo Add troubleshooting instructions to team lead page. -Met with Dr.Grasso to discuss troubleshooting for finding files. -Sanity check on files in CS Spanish Transcription Status for Whendy. Check if audios are clipped and ran through Whisper. -BIOBS012 → Clipped + Whisper Output -BILP035 → Clipped + Whisper Output -BIOBS011 → Clipped + Whisper Output -BIOBS009 → Clipped + Whisper Output -SSOBS012 → Clipped + Whisper Output -BILP013 → Clipped + Whisper Output -SSSE009 → Picnic + Important Event not administered. Cat Rescue located. -SSLP010 → Picnic not administered. Important Event and Cat located. -BILP010_Obs2 → Smartsheet indicates tasks were clipped but cannot be located *** -BILP020_Obs2 → Tasks not administered -SSOBS008--> No picnic or Cat administered. Clipped but could not locate files. -BILP029--> Files located. No cat. -SSLP008 → No picnic or important event. Cat rescue whisper located. -SSLP006 → Only important event. Picnic not ran through Whisper yet. -SSLP004--> Tasks not administered. -BIML001 → No picnic. Cat Rescue + Important Event whisper output located. -BIML002 → Smartsheet row blank.
Date Priorities Notes -Sanity check on files in CS Spanish Transcription Status for Whendy. Check if audios are clipped and ran through Whisper. -BIOBS012 → Clipped + Whisper Output -BILP035 → Clipped + Whisper Output -BIOBS011 → Clipped + Whisper Output -BIOBS009 → Clipped + Whisper Output -SSOBS012 → Clipped + Whisper Output -BILP013 → Clipped + Whisper Output -SSSE009 → Picnic + Important Event not administered. Cat Rescue located. -SSLP010 → Picnic not administered. Important Event and Cat located. -BILP010_Obs2 → Smartsheet indicates tasks were clipped but cannot be located *** -BILP020_Obs2 → Tasks not administered -SSOBS008--> No picnic or Cat administered. Clipped but could not locate files. -BILP029--> Files located. No cat. -SSLP008 → No picnic or important event. Cat rescue whisper located. -SSLP006 → Only important event. Picnic not ran through Whisper yet. -SSLP004--> Tasks not administered. -BIML001 → No picnic. Cat Rescue + Important Event whisper output located. -BIML002 → Smartsheet row blank. -Assws indo doe rh tracking of pprojectic specific look at folder structure and make sure we are satisfied with it there was no addiitional guidance to the transcribers -Update the section that includes the instructions for thre transcibers to reflect what was actually done/ some stuff is there but some stuff is not exactyly relevant On the wiki also be sure to enumerate when the column that indicates transcription status gets updated in this filer there is a folder that does x had to indicate who was to retranscribe and perofmr consensus. based off their values some got assigned consneus other -Continue Clipping -Finished clipping BISD024_Pre and began to clip BILP036_Pre -Sanity check on the timepoint days apart on the Shiny app -Confluence export PDFS of the entire connected speech SOP in the case Dr.Grasso asks Sanity Check on Parent Folder: -Catalan BILP020 and BISD011 Pre copied from project specific folder to parent. -Spanish BISE013 12m and BISE011 Obs1 copied from project specific folder to parent. From Shiny App: BILP014_Pre Span → Pre to 1 year follow up was <100 days. Changed year for 12m to 2023 BILP014_Pre Cat → Pre to 1 year follow up = 101 days. Changed year for 12m to 2023. BISD013_Pre Cat to Post was > 200 days. Didn’t see anything inherently wrong with the dates. Left as is. Seems pre and mid were also scheduled far apart. BILP025_Pre Span to Post was >400 days. Changed for 12m to 2024 from 2025. BIML002_Pre Span to Post was > 400 days.. Changed year from pre from 2021 to 2022. BISD013_Pre Cat to 6mo was > 400 days. Didn’t see anything inherently wrong with the dates. Left as is. BILP008_Pre Span to 6mo was > 600 days. Changed 6m year to 2021 from 2022. From Meeting with Dr.Grasso: Download the R desktop app to laptop app Install certain toolboxes double click shiny app click run app counts for sonia’s project Ranges 200-300 for 6m Ranges pre to 12m are going to be around the 300 range Ranges for pre to post are around the 50-70s -Verify the dates How would I verify the dates are correct for the time point visits? BILP014 Catalan example where the pre to post was over 400 days! Go to smartsheet > connected speech data analysis Catalan parent sheet. Madr information participant sheet CONNECTED SPEECH R dashboard -Clip and run Whisper -Finished clipping audio for BISE019_12m in Catalan. -Finished clipping the 3 parts for BILP027_12m in Catalan. -All changes are reflected on the smartsheet and Box. -I was not able to get access to TACC but will continue trying throughout the week. -Finished clipping: BILP028_6m and I am nearly done with BISE019_12m -Continue clipping samples -Run Whisper if there are nodes available -Finished clipping BISD019_6m -RAN WHISPER -Run Whisper for the Picnic Samples -Sanity check on transcriptions -Continue clipping priority samples Finished clipping BILP034_6m in Catalan. Finished clipping BISE022_6m in Catalan. Checked status of transcriptions in original folder. There are still a few missing: -Catalan: BISD011 Pre, BILP020 Obs1 -Spanish: BISE013 12m, BISE011 Obs1, As for whisper, I waited for a node my entire shift and was not able to get in. I have passed off the laptop to Aaliyah. -Continue clipping samples -Run Whisper Clipped BISE026_Pre in Spanish and BILP036_Pre in Catalan -Finished clipping samples on Audacity for BISE015 and BILP028 BILP028. part 2 → microphone tap at 5:57 ish also there is continuous buzzing throughout all the samples for this patient Part 2 teams notification at 10:10 -Meeting with Dr.Grasso -Run Whisper for Picnic Scene Samples. Then for other tasks. -Continue clipping audios I had started if time permits.
Date Priorities Notes -Finished clipping samples on Audacity for BISE015 and BILP028 BILP028. part 2 → microphone tap at 5:57 ish also there is continuous buzzing throughout all the samples for this patient Part 2 teams notification at 10:10 -Run Whisper -Continue Clipping -I attempted to run Whisper today for different samples but there weren’t any nodes available. I began to clip samples on Audacity instead but also had trouble with the headphones not connecting. -Clip samples and run whisper for recently clipped tasks. -Finished clipping for BISE022_6m -Finished clipping for BILP034_6m, pending clarification for MAINCat task. I’ll upload the tasks on box once clarified. For now I have saved the Audacity project under my notebook. -Finished clipping BISD019_6m -I was not able to get a node for Whisper so I will work on those next week as they’ve begun to accumulate. -Update reliability spreadsheet with consensus -Continue editing Whisper Catalan Smartsheet -Finished clipping BILP036_Pre -Started to clip BIE015_Pre -Finished updating the CS reliability sheet for Catalan and Spanish. Most congruent transcriber pair also updated on the sheet. -Finished editing the Catalan Smartsheet: BISE012 →Psychiatric case manifested as a neurodegenerative disorder. Patient not seen anymore. DELETE EVERYWHERE. BISE011 POST → Fixed. Note imp event clip could not be located according to smartsheet comment. BISE010 POST → Fixed with Dr.Grasso BISE010 3M → Picnic scene has not been clipped. BISE010 6M → Picnic scene has not been clipped. BISE007 12M → Fixed. Note imp event clip could not be located according to smartsheet comment. BILP012 PRE → Fixed. BIML001 12M → Fixed. Important event not administered. BILP019 3M → Fixed. Note imp event clip could not be located according to smartsheet comment. BILP018 PRE → Fixed. Note imp event clip could not be located according to smartsheet comment. BILP018 12M → Fixed. Note imp event clip could not be located according to smartsheet comment. BILP017 12M → Fixed. Note imp event clip could not be located according to smartsheet comment. BILP015 3M → Fixed. Note imp event clip could not be located according to smartsheet comment. BILP015 12M → Fixed. Note imp event clip could not be located according to smartsheet comment. BILP014 12M → Fixed. Note imp event clip could not be located according to smartsheet comment. BILP014 3M → Fixed. Note imp event clip could not be located according to smartsheet comment. BILP013 12M -> Fixed. Note imp event clip could not be located according to smartsheet comment. BILP007 PRE → Fixed. Picnic scene not administered. -Update the reliability spreadsheet with new rely values -Check that transcribers have followed instructions of replacing the pre-consensus transcript with the post-consensus -Go into Catalan Whisper Smartsheet and ensure that all Pre-R01 tasks have been marked as “completed” (if applicable) Transcribers have began to populate the Box folder with the updated transcriptions (post-consensus) in the original folder. They have been updating the folders within Longitudinal PPA Project with the updated transcriptions and replacing the (pre-consensus) transcriptions. Note that Pre-R01 Tasks consist of: Cat Rescue, Picnic, and Important Event -Confirm participant timepoints that could not be found on the smartsheet -Run RELY for latest re-transcription sample -I ran RELY for BILP017 Pair 2 +3 (JC + HM) → include consensus met in smartsheet % Utterances: 100 % Words: 93.0 Pair 1+3 (NMC + HM) % Utterances: 73.3 % Words: 91.0 -Finished clipping BILP035_Post and BISD024_Pre Minor tweaks to Connected Speech Documentation: After the consensus meeting: Update the original transcript that is located in its original folder (e.g., B--Connected Speech_Data) based on the agreed changes. Delete any old or unreliable transcript versions from the original folder so only the final version remains. The updated transcript is now the final, reliable version and can be used for analysis.  Record all reliability calculations and decisions in the Connected Speech Reliability Smartsheet. Update Transcriber 1 and Transcriber 2 in the Smartsheet with the final reliability values (the most matching pair). Previous reliability results will be documented in the project specific consensus folder within the transcriber spreadsheet. -Finish Running PRE-RO1 Catalan samples that crashed + update the smart sheet. Noticed some participants could not be found on the smart sheet . To find these go to the Catalan parent sheet and check if all assigned tasks have been marked as “completed” as they may not be updated. If not, take note of the patient timepoints that are actually missing. BISE012 → DELETE EVERYWHERE BISE011 POST BISE010 POST, 3M, 6M. BISE007 12M BILP012 PRE BIML001 12M BILP019 3M BILP018 PRE, 12M BILP017 12M BILP015 3M, 12M BILP014 12M, 3M BILP013 12M BILP007 PRE -Prioritize clipping for Catalan to begin getting the samples ready for the transcribers! -It looks like most consensus meetings were done yesterday and Saturday! -Helena + Jaume and Helena + Whendy are nearly done with consensus. Pending 7 more! -Run RELY for updated transcriptions on samples -It looks like most consensus meetings were done yesterday and Saturday! -Helena + Jaume and Helena + Whendy are nearly done with consensus. Pending 7 more! -Prioritize clipping Pre, 12m, 6m, NOT Post or Mid! → Filter Clipping Smartsheet to show only those AND any that are not completed!! -Finish running audio for SSSE10 -Check for any updates on transcriber spreadsheet -Run Cat Rescue Catalan samples through Whisper -No new updates on the transcriber’s spreadsheet. -Finished Clipping SSSE010_Post -Continued to clip BILP035_Post -I wasn’t able to run Whisper because I couldn’t get a TACC token, but I’ll be sure to do so on Monday so I can finish the Catalan samples. -Run RELY with transcriber 3 samples Ran RELY for BILP010_Pre AEQ (Transcriber 1) + HM (Transcriber 3) % utterances: 82.4 % words: 75.2 WAM (Transcriber 2) + HM (Transcriber 3) % utterances: 87.5 % words: 74.2 Ran RELY for BISE013_12m MVR (Transcriber 1) + HM (Transcriber 3) % utterances: 72.7 % words: 74.5 WAM (Transcriber 2) + HM (Transcriber 3) % utterances: 100 % words: 89.8 -Continue to run Whisper -Check updates from consensus -Continue to clip audios -Finished running all samples for the Cat Rescue Spanish samples! -Pending clarification for BISE011 audio labeling -Run Whisper for samples -Finished running Whisper for the following Spanish samples Pre, 3m, 6m, 12m, Post for all except BILP014 BILP006 BILP007 BILP008 BILP009 BILP010 BILP011 BILP012 BILP013 BILP014 (12m only) I had issues with the Whisper laptop. It was running very slow and kept closing tabs. At some point Whisper stopped generating the transcripts but I was able to finish the samples. All changes are reflected on Box and Smartsheet. I also spoke with Arely about a few questions I had regarding clipping and using Audacity. Thanks!
Project 1: Overview of Connected Speech Project Date Priorities Notes -Finish running Whisper for samples I have clipped -Begin running Whisper for samples Aaliyah assigned. WFH Today Catalan Bridge OBS Samples (All completed): BISE010:Obs1 (Audios 1-9) and Obs2 (1 audio) BISE011: Obs1 (Audios 1-11) BILP013: Obs1 (Audios 1-9) and Obs2 (1 audio) BILP010: Obs1 (Audios 1-8) BISE006: Obs1 (Audios 1-9) and Obs2 (Audios 1-9) Spanish (recently clipped) all audio samples ran through Whisper for: SSLP009 BILP030 BILP033 BISD023 BIML004 I will continue to run the pending samples Aaliyah assigned me next week. -Finish Clipping BISD023_Post -Run through Whisper completed Span. Samples -Finished clipping BISD023. All changes are reflected on the Smartsheet and Box! I also began writing Whisper for S1+S3 Spanish Tasks. Picnic: BILP030_6m BILP033_6m SSLP009_6m BISD023_Post Important Event: BILP030_6m BILP033_6m SSLP009_6m BISD023_Post BIML004_Pre Teeth: SSLP009_6m BISD023_Post BIML004_Pre BILP030_6m BILP033_6m MAIN DOG: BIML004_Pre BILP030_6m BILP033_6m SSLP009_6m Frog Where: BIML004_Pre BILP030_6m BILP033_6m SSLP009_6m -Finish clipping BILP033_6m -Continue to clip WAB samples that contain the picnic scene From meeting with Dr.Grasso: -Refer to the parent Smartsheet to see which participants have not had the Picnic Scene WAB sample clipped! Those are of highest priority for the project! -Finished clipping BILP033 -Halfway done clipping BISD023 -Begin clipping WAB samples that have the picnic scene -Meet with Arely to discuss clipping training/feedback I clipped BILP030_6m. All changes reflected on Smartsheet and box! Halfway done clipping BILP033_6m Note: The Audacity project file can be found in Box lab notebook in order to continue on different desktop. Left off at 22 m and 32 s. No interruptions yet. Received feedback from Arely regarding my clipping training. Green light to clip samples. -Begin to clip new audio samples -Clipped BIML004. All tasks except Bridge and Picnic were administered. All outputs were put into their corresponding folders in Box. -Completed clipping SSSE010_Post. Need to export the file and add everything to its designated folders. -Edit Audacity project and upload to Lab Notebook on Box -Finished clipping project and made changes to the clipping. Such as removing accidental loops I made, etc. All files, including the Audacity project, have been uploaded to my lab notebook on Box under “Clipping Practice” folder. -Questions: How to determine which WAB samples are missing from the longitudinal project? Were the samples underneath green row on CS Reliability sheets ever confirmed with Sonia? -Meet with Dr.Grasso + Jada to discuss next steps for project Clipping: Meeting notes: Make the following changes to Smartsheet conditional formatting below 99.5 highlights if below .0795 to account for running -Add automation workflow to the smartsheet -Finish clipping practice WFH Added conditional formatting to columns on the smartsheet. Also made sure the %s rounded to whole numbers. Completed the clipping practice. Made sure the initials of all transcribers were present in the file names for both Catalan and Spanish in the Longitudinal PPA project. -Finish running reliability for Catalan samples Ran reliability for all Cat samples: BILP012 BILP017 BILP020 BISE004 BISE010 BISD011 BISD013 BISD014 BILP020 BILP027 BISD015 BISD008 BILP008 BILP009 BILP010 BILP012 BILP014 BISE006 BISE007 BISD007 -Finish running reliability for Spanish samples. -Copy Rater 1 CATALAN Files to Longitudinal speech project folder Finished running reliability for the Spanish samples: Note: BISD015 had 100% in both words and utterances. Unsure if this is normal. Began to copy over rater 1 CATALAN files into respective folder. Pending adding initials of transcriber into the file name. I finished running these samples through Clan and have updated the smartsheet to reflect the results. The sheet seemed to have settings that highlighted any changes in yellow. BISE013 BISD011 BISD013 BISE016 BISE011 BILP009 BILP011 BILP013 BILP014 BISE011 BISE013 BILP022 BISE014 BISD012 BILP011 BILP018 -Finish copying over Rater 1 Files COPIED: -BISE011 Obs1 -BISE013, 12mo -BILP022, 12mo -BILP018, Pre -BILP011, Pre -BISE014, Pre -BISD012, 12mo -BISD015, 12mo Note: -BILP007, Pre → Transcriber 1 file could not be located. Transcriber 2 not located, uploaded by Sonia. -BILP008, Pre → Transcriber 1 file could not be located. Transcriber 2 located. -BILP009, Pre → Transcriber 1 file could not be located. Transcriber 2 located. -BILP010, Pre → Transcriber 1 file could not be located. Transcriber 2 located. -Meet with Dr.Grasso to discuss next steps. Met with Dr.Grasso regarding next steps for the longitudinal PPA project. Beginning to work on: @Estrella Palomo to copy all Rater 1 files over to box Reliability folder for the Longitudinal PPA project https://utexas.app.box.com/folder/340279736665?s=zzqulcftyamfs5pid1yu6ww89bbzm9kf by Files copied to Transcriber 1: -BISE010, Pre -BISE013, Pre -BISE011, Pre -BISE016, Pre -BISD013, Pre -BISD011, Pre -BILP006. 12mo -BILP009, 12mo -BILP 011, 12mo -BILP013, 12mo -BILP014, 12mo Note: -BISE005, Pre → Transcriber 1 File could not be located. Transcriber 2 located. To-Do: -2 BRIDGE Files -Obs -Pre-Ro1 Catalan → Double check Jimena’s smartsheet input on box -Fill out extra boxes on smartsheet -Catalan Cat Rescue Picture Story Description Pre-R01 Notes: ENSURE ALL ARE IN THEIR CORRESPONDING FOLLOWERS Clinical: 12m, 6m, pre, post, and MID Eval: EB1: Brushing teeth, important event, picnic scene Observational: OBS 1 and 2 PRE-RO1: Pre, Post, 12m, 6m, AND 3m Ran whisper for: BISE024 (Pre) BISE025 (Pre) BISE 022 (Pre) BISE021 (Pre) BISE013 (12m) BISD023 (Pre) BISD019 (Pre) BISD018 (Pre) BISD016 (Pre) BISD016 (12m) No sample 6 BISD014 (Pre) BISD014 (12m) BISD 013 (Pre) No sample 9 BISD013 (12m) BILP034 (Pre) BILP029 (Pre) No sample 2 BILP028 (Pre) BILP024 (Pre) No sample 6 or 8
Tasks Completed: Moved incorrectly filed Cat samples to their correct folder in LongitudinalPPAPicnic_Project_Reliability Located “BILP008_12mu_PicnicDescription_CAT” in box within 101 Connected Speech Data as instructed in the 7.Connected Speech/Transcription Reliability. Copied it into the longitudinal PPA folder. Tasks Completed: Practice going through 6.Acoustic Derivation Guide on Whisper. Done successfully. Downloaded CLAN to both personal and lab laptop. Notes on Acoustic Practice: Samples can be found in Box Lab Notebook. This was the final file with all acoustic derivations HERE Need to practice running more files as I had a bit of trouble with the coding. Mostly due to spacing and file names. Acoustic Practice Excel Sheet: Worked on reading through 6. Acoustic Derivation Guide in preparation for next week’s meeting with Dr.Grasso. I was unable to run Whisper as tokens were unavailable I have finished filling in the column on Pilar's sheet. The file total (172) on the spreadsheet matches what is in the Box folder. The only pending files are the ones previously mentioned. To summarize:   -63601 & 537728 were missing the picnic scene in the FTLD folders -There was a patient that I believe to be misnumbered. The spreadsheet says "1748597" but based on the CS administration date I believe it is meant to be "1748957" -BILP006 had two picnic scenes in the folder. The file names are identical. Same for: 285086, 908807, 1705735 -Finally, SSOBS002 had the incorrect CS administration date on the spreadsheet. I did not edit it but made a comment.   Notes from Dr.Grasso: 63601 & 537728 were missing the picnic scene in the FTLD folders There is no ID 63601 in the sheet. I believe this is a typo There was a patient that I believe to be misnumbered. The spreadsheet says "1748597" but based on the CS administration date I believe it is meant to be "1748957" Yes it is, I copied over the correct sample to the folder BILP006 had two picnic scenes in the folder. I am not sure where you saw the two scenes. Let's look together  I moved over all the relevant BILP006 files because I am guessing you did not find them in Pilar's folder   The file names are identical. Same for: 285086, 908807, 1705735 I moved over the files for these three individuals as well 285086, 908807, 1705735, so this needs updating in the sheet but I cannot locate the version you said you updated Finally, SSOBS002 had the incorrect CS administration date on the spreadsheet. I did not edit it but made a comment.  Because I cannot locate the spreadsheet edits you made, we need to look at this together next week. I changed it by one day if that's the error you saw (7/3=> 7/2) and also deleted the entry from a few days prior because I could not find any folder or file indicating that samples were collected on that day (6/28).   Worked on copying picnic scenes for SpeechFTLD_A. All changes and comments are reflected on the spreadsheet. Comments made: Patient 1748597 not found in the SpeechFTLD_A folder but there was a patient 1748957 that matched the CS administration date. BILP006 had two picnic scenes in the folder. Two patients had picnic scene missing (663601 & 537728) Completed the copying of picnic scene for SpeechFTLD_B. All changes reflected on the spreadsheet. Begin reading new paper I was still not able to run Whisper. Laptop was not able to be found and queues were not available. Assisted Dr.Grasso with filling in missing information for Dr.Santos' spreadsheet Create 1password account Finish Whisper Runs for BISE013 Finish Whisper Runs for BILP 035 Give heads-up about schedule change Was not able to run whisper samples. Troubleshooted with Arely but there were no nodes available even after trying all the different gpus. Spoke with Aaliyah about next steps for whisper once I have completed the samples that are in progress. Began to read new paper “Screening for early Alzheimer’s..” Also went over the general workflow with Arely. Read Andrew’s Paper Meet with Arely Finish Whisper Runs Update Smartsheet -Read and took notes on Andrew’s paper. -Met with Arely to discuss Whisper process and learn to better navigate Box and Smartsheet. Comments were left on BILP035 and BISE013 regarding progress on samples. -Next papers to read: ”Primary progressive aphasia and the evolving neurology of the language network” Screening for early Alzheimer’s disease: enhancing diagnosis with linguistic features and biomarkers Was not able to run Whisper for samples BILP035 and BISE013 as there weren’t nodes available. Waited about ~1 hour to see if any became available. I will resume progress and finish running all samples for these patients on . Ran whisper sample for BISD012 Practiced running 5 bridge samples. Four of the samples were from BILP035 and one from BISE013 Sample for BISD012 was successfully ran through Whisper. The samples ran were. Completion is reflected on the smartsheet. Sample for BISE013_BACC001_Bridge_Cat_Obs2_20251008.wav was completed successfully. Change reflected on box. Pending updating the smartsheet to reflect change. Samples 1,3,4 and 6 for BILP035 were successfully ran through Whisper. I am pending to update the smartsheet. Box reflects the changes made. I ran these all at once given as they were the same patient.
Practiced running Whisper with individual speech samples. Downloaded UT VPN I was able to generate a successful transcript output on whisper following the instructions on Also installed the VPN Need to practice running multiple related samples at once Get used to naming conventions used on Box Note to self: When making folder, ensure it is under “Whisper Runs” and not just in the general hub. Naming convention: speaker_date_topic.wav @Estrella Palomo download VPN @Estrella Palomo to practice running Whisper on samples in lab notebook: . See Connected Speech structure for Whisper Guide by @Estrella Palomo to read and take notes of paper Dr. Grasso sent by @Estrella Palomo to review folder structure on Box and accompanying Smartsheet to see how we can tell which clipped samples need to be run through Whisper by : @Estrella Palomo to Review connected speech documentation @Estrella Palomo Meet with Arely and Aaliyah to go over clipping task structure and running Whisper and learn how she manages that

Research Paper Notes

Paper What they wanted to find out How they did it What they found Why it matters for our project Gorno-Tempini et al. (2011), Neurology — "Classification of primary progressive aphasia and its variants" Create one agreed-upon set of rules doctors everywhere could use to diagnose the three types of PPA A panel of experts agreed on official guidelines Not a study of actual patients They defined the 3 types of variants: Nonfluent → grammar and speech-prod. problems Semantic → loss of word meaning Logopenic → word finding + rep problems Use same 3 levels to lock down diagnosis: -Symptoms -Pathology/Genetics -Brain scan Note: Each type tends to involve a certain brain region and disease, but only as a tendency, not a guarantee. @Estrella Palomo to review the videos here and take notes on what she sees with respect to speech and language symptoms present in each variant. Add notes by participant code by Foundational study! Used to define the 3 variants. PPA studies classify patients using these rules. Hardy et al. (2024), European Journal of Neurology — "Symptom-based staging for logopenic variant primary progressive aphasia" A good follow-up including more variants Map out how the logopenic type progresses over time, from earliest signs to most severe, based on what families actually observe Surveyed caregivers twice about how symptoms appeared and worsened. 2 Phases: i. Exploratory Survey -34 caregivers -order symptoms based on Reisberg scale ii. Consolidation Survey -37 caregivers in the UK and Australia -rearrange draft of 6 stage framework They identified 6 stages ranging from “Stage 1: mild” to “Stage 6: Profound”. Early on there was trouble finding words, hearing, some memory/navigation problems. Much later was the development of being unable to understand others, difficult speech to understand, full dependence… Also flagged “milestone” symptoms which indicate big life changes. Some examples are having to stop working, driving, etc. Note: Stage durations varied from 6 months to 8 years. Shows one way to chart how the disease unfolds over time. Alternative approach to longitudinal study by studying caregivers. Note: Authors say the stages were built around English speakers and are unlikely to carry over to other languages. Canu et al. (2025), Neurology — "Connected Speech Alterations and Progression in Patients With PPA Variants" Figure out the best way to tell the three types apart, and watch how each type's speech changes over time Studied 95 patients. Each described the Picnic Scene to study natural speech, took standard language tests (naming, etc.), and most also had brain scans done. A smaller group came back about 10 months later. The semantic variant was the easiest to identify with the basic language tests. Nonfluent and logopenic difficult to tell apart with basic tests. Connected speech examinations and brain scans were the best at distinguishing between the two. Notes on individual variant declines: @Estrella Palomo to add some notes about how each of these measures was defined in the study by and to elaborate as to what they found at time 1 and also the changes over time specifically for connected speech. What measures did they include and which were significant nfvPPA: sound errors, svPPA: more meaning, naming, and grammar errors lvPPA: slower speech and shorter sentences Shows that analyzing natural speech is a good way to measure language decline over time. Strong support for how we can track decline across mono and bi patients. Ash et al. (2019), Brain & Language — "A longitudinal study of speech production in primary progressive aphasia and behavioral variant frontotemporal dementia" Track how patients' speech changes over time by identifying patterns of decline and whether the decline matched up with brain shrinkage. 48 patients total, all 3 variants represented + bvFTD Described the Cookie Theft Picture at two visits 1 yr+ apart. They measured speech speed, sound errors (per 100 words), dependent clauses, and well-formed sentences. Everyone’s speech decline over time. The nonfluent variant was the one that declined the most (in ALL measures). Semantic variant simplified its grammar. Logopenic slower speech + grammatical errors -Speech decline did NOT line up with decline on standard memory/thinking tests. Shows that connected speech does a good job at identifying these issues in comparison to basic examinations. -Worse grammar = tissue loss in language areas of brain Similar template to project: Follows sample patients' natural speech across various visits. Shows that connected speech captures important + unique information about the decline of each variant. Also notes the small sample sizes and how this is a challenge in the field. Mueller et al. (2018), Journal of Clinical and Experimental Neuropsychology — "Connected Speech and Language in Mild Cognitive Impairment and Alzheimer's Disease: A Review of Picture Description Tasks" Note: this is about Alzheimer's/MCI, not PPA. I thought it would be an interesting citation for methods. Literature review of 36 studies on language of individuals with AD or MCI. Looked at all the research using picture description tasks (Picnic + Cookie Theft) Wanted to answer: what gets measured, does it work, and can it catch decline early? A review (no new patients) that looked at 36 existing studies, most using the Cookie Theft picture, and the rest Picnic Scene. Note they only looked at English based journals. 1,100 Alzheimer's patients and ~270 with mild cognitive impairment. Picture description separates patients with AD from healthy individuals. Most useful measure is semantic content. Patients tend to produce “empty speech” which is essentially many words with little ideas or meaning. Grammar also simplified sooner than assumed. Supports the Ash study in saying that connected speech examinations are a good way to holistically examine a patient’s speech patterns. Discusses value of connected speech analysis: -CS uses many mental processes at once -CS most closely approximates language production in everyday contexts than standardized tests. -CS provides quick means of assessment and low burden on participant Also notes the gaps in research: -Cross-language transfer -Lack of diversity -Lack of longitudinal studies Costa et al. (2019), Alzheimer Disease & Associated Disorders — "Bilingualism in primary progressive aphasia: a retrospective study on clinical and language characteristics" Focus: bilingual PPA, all three variants Describe the manifestation of PPA in bilingual individuals. Specifically, which language is affected first and a comparison of decline rate across languages. 33 bilingual PPA patients by looking back at records: -13 from previously published cases -20 from new cases (across 5 countries) Coded each person’s two languages to determine dominance, when the 2nd language was learned, proficiency, usage frequency. Note: did not have a monolingual group for comparison. Word-finding was the most common symptom (and the first) → mostly noticed in the 2nd language (but often still the dominant one). On average, both languages declined in parallel (supports sharing of language networks for bilingual individuals). Despite this, each variant still showed their differences in decline. Discussion regarding “language mediators” in L1 and L2 differences. Defined as proficiency, age learned, usage frequency, etc. Decline may be reflecting this instead of the disease itself. Supports basis of our project: -Calls for a longitudinal study with a monolingual vs bilingual comparison. -Emphasis on assessing patients in both languages and carefully measure “language mediators” in order to determine if a difference in language decline is due to a person’s language profile or the disease itself. Tippett et al. (2025), Journal of Alzheimer's Disease — "Written Picture Descriptions Distinguish Variants of Primary Progressive Aphasia" Whether written (not spoken) picture descriptions can tell the three PPA variants apart, and which analysis method works best in a clinic. 49 patients (16 lvPPA, 17 nfvPPA, 16 svPPA) + 18 healthy controls handwrote a Cookie Theft Picture description (no time limit). Analyzed 3 ways: i. Parts of speech: (counts/% of nouns, verbs, particles, etc.) via Open Brain AI, hand-checked ii. Content: Out of the 52 things usually pointed out about the image how many were mentioned by each person and how descriptive they were. iii. Checked how many of 26 common words (the ones most healthy participants use for the picture) each person included. Each variant had a distinct profile. nfvPPA: used fewer words and left out connecting words… similar manifestation that would show up in speech svPPA: wrote the least overall, named mostly single words and little phrases… fits the profile of not remembering definition of words lvPPA: wrote most simiarly to the healthy control group but would often go off-topic and use many “filler” words that wouldn’t have much substance to the task. To differentiate between the variants: counting the “expected ideas” (the 52 words most mentioned) was best for differentiating svPPA from lvPPA, and looking at word types was best for spotting nfvPPA Supports the idea that looking at natural speech (even in writing) picks up the specific differences between the three types that basic examinations miss. Paper includes list of “expected ideas” for reference de Leon et al. (2026), Aphasiology — "Reported symptoms and patterns of language impairment in bilingual speakers with primary progressive aphasia: a retrospective study" In bilingual people with PPA: what are the first symptoms families notice, which of the person's two languages gets hit first, and which one holds up better? And is that pattern better explained by which language was learned first (L1 vs L2) or by which language is the person's stronger/more-used one (their "dominant" language)? Looked back through the medical charts of 69 bilingual PPA patients (22 nfvPPA, 31 svPPA, 16 lvPPA) From each chart they pulled out the first symptoms reported by the patient/family, which language was affected first, which was less preserved, and background on the person's languages (when each was learned, which was stronger). Only the first visit was used, NOT a longitudinal study. Two reviewers did the coding and agreed 90% of the time. Most first symptoms matched the standard PPA picture, but a few were unique to bilinguals. For example trouble switching between languages or trouble translating. MAIN FINDING: the weaker/less-used language was usually the first to be affected (93% of cases) and the less preserved one (88%), no matter whether it happened to be the first or second language learned. One exception: in lvPPA, the second-learned language was always affected first. And svPPA more often showed both languages slipping equally supports focusing on language dominance (not just first vs. second language) when we look at how bilingual patients decline. limitations: it relied on families' memories rather than direct testing, cross-sectional, and everyone was tested only in English.
Notes on Screening for early Alzheimer’s disease Notes on Andrew’s Paper Issues: standard assessments lack validity and don’t accurately reflect daily communication. there has been more focus on english-speaking populations which can contribute to poor diagnosis of spanish-speaking populations Addressing it: Connected speech analysis. Based on a previous study, it has been shown to show linguistic features and also underlying neurodegenerative patterns. “Verb Paradigm”- different forms a verb can take. In Spanish there are about 50 while in English there are ~3-7. An English assessment being used on a Spanish speaker would fail to to provide a good description of the patient’s performance as it is not readily made to assess the difference in language nuance. For example, grammatical gender. Paper study: Analyzed connected speech features from connected speech data of Spanish speakers with PPA → to find linguistic markers that can effectively distinguish between the 3 clinical variants and healthy controls. morphosyntactic behaviors- ways we combine meaningful word parts (morphemes) and words (syntax) to build sentences. This reflects a language’s rules for word structure, order, and grammatical agreement. Goals: 1.) Identify clinical + linguistic variables that are sensitive to a specific variant pattern of language breakdown 2.) evaluate new morphosyntactic of Spanish (such as grammatical gender) to identify linguistic markers of impairment that go beyond what is available in English. Overall Goal: Because of our unclear understanding of how Spanish-specific features manifest across the three variants, this study sought out to find whether they can complement existing diagnostic markers and improve the classification of PPA in Spanish. Methods: 67 participants: 19 nfvPPA, 22lvPPA, 15 svPPA, and 11 controls. All participants were native Spanish speakers from Spain. They underwent clinical evaluations that included memory, language, executive function, and visuospatial skill assessments. All participants met current consensus for PPA diagnosis. Results: All PPA groups performed under control in neuropsychological testing. The largest impairments were in language dependent measures such as semantic/phonemic fluency, confrontation naming, and verbal memory. nfvPPA showed preserved naming and semantic knowledge but presented agrammatism and motor speech deficits lvPPA showed impairments in naming and working memory. svPPA showed deficits in naming All variants showed intact visuospatial and basic perceptual abilities Definitions: semantic/phonemic fluency: Phonemic is letter fluency. For example, “generate words beginning with a specific letter.” Semantic is category fluency. For example, “generate words from a specific category.” confrontation naming: assess and individual’s ability to retrieve words for viewed objects verbal memory: assessments designed to evaluate memory and learning through tasks that involve the repetition and recall of words. Connected speech task: All participants were asked to describe the picnic scene from the Western Aphasia Battery Revised Results for lexico-semantic processing and anomia: both nfvPPA and lvPPA produced more repetitions than controls. lvPPA also produced more retracings than svPPA and controls. lvPPA had higher frequency nouns and verbs relative to all other groups. Word correctness: individuals with nfvPPA produced more concrete words than all other groups and control > lvPPA Type to token Ratio (TTR): higher in the nfvPPA groups compared to lvPPA, svPPA, and controls. This indicates a more varied vocabulary despite reduced fluency. Results for Traditional morphosyntactic measures: nfvPPA produced the lowest utterances compared to the other groups. Number of relative pronouns was also the lowest. lvPPA, svPPA, and controls showed greater use of embedded clauses. Verb to noun ratio: lvPPA and svPPA produced more verbs than the nfvPPA and controls. Mean of length utterances higher in lvPPA, svPPA, and controls higher than those with nfvPPA. Discussion: Patterns observed in the study are consistent with previous samples from English speaking populations. The nfvPPA group demonstrated a marker reduction of utterances which has been documented in both English and Spanish speakers with nfvPPA. Lexical and fluency disruptions (repetitions and retracing) in lvPPA also reflect previous work in Span and Eng. svPPA reduced concreteness and lexical diversity also aligned with previous studies. The results reinforce the profiles for each variant in English and Span which suggests a cross-linguistic pattern Notable pattern: absence of frequency effects in the svPPA group Notes on: “Cognitive-linguistic skills in production of expository discourse: Insights from longitudinal changes and neural correlates in primary progressive aphasia” Aim of study: The study aims to see whether difficulties with grammar/sentence complexity in PPA are explained only by damage to language regions or also by damage to brain regions involved in attention/other cognitive skills Discourse = connected speech that goes beyond single words or sentences. (E.g. describing a picture, telling a story, explaining an idea or opinion). Analysis Microstructural- vocabulary, morphology (word endings like -ed -s), syntax (sentence structure) Macrostructural- how sentences connect: cohesion (link btwn sentences) and coherence (overall flow) Producing discourse requires: language skills, attention, working memory = discourse is a sensititive marker for neurodegeneration PPA: Neurodegenerative condition → causes progressive language decline Language is affected first → cognitive skills (attention, memory, executive function) decline later Early PPA affects classic left perisylvian language regions = group of brain regions in the left hemisphere that sit around sylvian fissure Expository Discourse (used in this study) → describing a picture or scene, no time sequence, focuses on sentence/grammar structure. Cookie Theft Picture. Participants: 90 individuals with PPA 27 nfvPPA 30 lvPPA 33 svPPA 36 healthy controls Typical Patterns across PPA variants: nfvPPA- short sentences, simple grammar lvPPA- word finding difficulty, phonological errors svPPA- grammar mostly intact, reduced content word use Neuroimaging nfvPPA- brain thinning in left frontal areas involved in speech production and planning sentences lvPPA- left temporoparietal and temporal areas = pathway important for word retrieval, sentence construction svPPA- temporal lobes- word meaning, concepts. includes both left and right temporal areas. Researchers looked at: Utterance length (how long are the sentences?) Sentence complexity (use of embedded clauses) Flawed sentences (grammatical error) Idea density (how much information is packed into speech) Semantic → little connection between brain damage and grammar Nonfluent and Logopenic → linked to both language areas and attention areas