Albanian Language Expert – AI Data Annotation
iMerit Technology
/job-details/11192
Remote | Freelance | Flexible Hours
iMerit Scholars brings together skilled professionals from around the world to help build safer, more accurate and more culturally aware AI. Our experts work on real AI projects for leading technology companies, shaping how AI systems understand and respond in their language.
About the Project
We're looking for Fluent Albanian speakers to help build privacy-safe training data for AI systems. You'll identify and tag personally identifiable information (PII), such as names, addresses, phone numbers and ID numbers, across Albanian text, audio recordings and transcripts, and label visual content in images using bounding boxes. This is fully remote, flexible work that you can fit around your existing commitments.
What You'll Do
- Read Albanian text and tag all personally identifiable information (names, addresses, phone numbers, email addresses, ID and account numbers, dates of birth and similar details) according to project guidelines
- Listen to raw Albanian audio recordings and mark the segments where PII is spoken, with accurate timestamps
- Review Albanian audio transcript snippets and tag PII within the transcribed text, checking against the audio where needed
- Draw precise bounding boxes around specified objects or regions in images (for example, visible text, documents, faces or other PII) and apply the correct labels
- Apply tagging categories consistently and handle ambiguous or edge cases as defined in the guidelines
- Flag unclear audio, poor-quality images or content that falls outside the guidelines
What We're Looking For
- Fluency in Albanian, with strong reading and listening comprehension, including natural, informal and fast-paced speech
- Strong English reading and writing skills (B2 level or above), as guidelines and instructions are in English
- Excellent attention to detail and the ability to follow detailed tagging guidelines consistently
- Comfort working with annotation tools for text, audio and images
- An understanding of what counts as personal or sensitive information, and respect for data confidentiality
- A reliable computer, good-quality headphones and a stable internet connection
Nice to Have
- Familiarity with regional variation across Albanian-speaking communities (e.g. Gheg and Tosk dialects, usage in Albania, Kosovo and North Macedonia)
- Previous experience with data annotation, PII/data redaction, transcription or bounding-box labelling
- Background in transcription, translation, proofreading, data entry or content moderation
- The chance to help build safer, privacy-respecting AI for Albanian speakers
- Access to future projects on the iMerit Scholars platform
iMerit Scholars brings together skilled professionals from around the world to help build safer, more accurate and more culturally aware AI. Our experts work on real AI projects for leading technology companies, shaping how AI systems understand and respond in their language.
About the Project
We're looking for Fluent Albanian speakers to help build privacy-safe training data for AI systems. You'll identify and tag personally identifiable information (PII), such as names, addresses, phone numbers and ID numbers, across Albanian text, audio recordings and transcripts, and label visual content in images using bounding boxes. This is fully remote, flexible work that you can fit around your existing commitments.
What You'll Do
- Read Albanian text and tag all personally identifiable information (names, addresses, phone numbers, email addresses, ID and account numbers, dates of birth and similar details) according to project guidelines
- Listen to raw Albanian audio recordings and mark the segments where PII is spoken, with accurate timestamps
- Review Albanian audio transcript snippets and tag PII within the transcribed text, checking against the audio where needed
- Draw precise bounding boxes around specified objects or regions in images (for example, visible text, documents, faces or other PII) and apply the correct labels
- Apply tagging categories consistently and handle ambiguous or edge cases as defined in the guidelines
- Flag unclear audio, poor-quality images or content that falls outside the guidelines
What We're Looking For
- Fluency in Albanian, with strong reading and listening comprehension, including natural, informal and fast-paced speech
- Strong English reading and writing skills (B2 level or above), as guidelines and instructions are in English
- Excellent attention to detail and the ability to follow detailed tagging guidelines consistently
- Comfort working with annotation tools for text, audio and images
- An understanding of what counts as personal or sensitive information, and respect for data confidentiality
- A reliable computer, good-quality headphones and a stable internet connection
Nice to Have
- Familiarity with regional variation across Albanian-speaking communities (e.g. Gheg and Tosk dialects, usage in Albania, Kosovo and North Macedonia)
- Previous experience with data annotation, PII/data redaction, transcription or bounding-box labelling
- Background in transcription, translation, proofreading, data entry or content moderation
- The chance to help build safer, privacy-respecting AI for Albanian speakers
- Access to future projects on the iMerit Scholars platform