La Moakhza: Apologies in Egyptian Arabic
A comparative study of apology strategies produced by native speakers and an LLM-based chatbot in two socially consequential scenarios.
Full-paper summary
The project compares how Egyptian Arabic speakers and ChatGPT manage responsibility, repair, explanation, and face in apology situations. It asks not only whether the model can produce recognizable apology formulas, but whether its pragmatic choices resemble the negotiation found in human interaction and whether it sustains the requested Egyptian variety.
The answer is mixed. The model is highly fluent at explicit apology and repair, but it is more uniformly conciliatory than the speakers. Human responses contain more explanation, resistance, blame, denial, and minimization. The model also shifts among Egyptian Arabic, Modern Standard Arabic, and literal English-influenced phrasing.
Design
The first situation concerned an employee who was late or absent; the second concerned responsibility after a car accident. Responses were segmented into apology strategies such as acknowledgment, explanation, apology, lack of intent, offer of repair, promise of forbearance, blame, denial, and minimization.
Main findings
- ChatGPT strongly favored accepting responsibility and offering concrete repair in both situations.
- Native-speaker responses relied more on contextual explanation and, in the accident scenario, on contesting responsibility through blame, denial, or minimization.
- The model produced substantially more strategy units overall, making its apologies elaborate but also formulaically over-accommodating.
- Dialect control was unstable: Egyptian Arabic alternated with Modern Standard Arabic and translated English-like expressions.
Results visualized
Limitations and significance
The study includes only six participants, two scenarios, and one model version. Strategy counts are descriptive and should not be generalized to all Egyptian Arabic speakers or LLMs. Future work should expand participant demographics, vary power and social distance, compare multiple systems, and have independent annotators assess both strategy and dialect naturalness.
The project nevertheless identifies a useful evaluation principle: culturally appropriate language technology should be tested for pragmatic calibration and variety consistency, not only for fluent text generation.