Pragmatic Interpretation of Irony in Moroccan Darija: A Comparative Study of Native Speakers and Large Language Models
Audio version created with Paper2Audio.
Listen on Paper2Audio
Pragmatic Interpretation of Irony in Moroccan Darija: A Comparative Study of Native Speakers and Large Language Models
1. Background of the Study
Pragmatic theory distinguishes what a speaker literally says from what a speaker means, and treats the gap between the two not as a marginal complication but as the ordinary condition of everyday communication. Recovering speaker meaning depends on inference: the hearer combines the linguistic signal with contextual assumptions, shared background knowledge, and beliefs about the speaker's communicative intention in order to arrive at an interpretation that the sentence itself does not encode. Irony represents one of the clearest instantiations of this inferential requirement. Rather than functioning as simple semantic inversion, as classical antiphrastic accounts once proposed, verbal irony is now widely understood within pragmatics as a mechanism through which a speaker echoes a proposition, expectation, or norm while simultaneously signalling a dissociative or critical attitude toward it. Correctly interpreting an ironic utterance therefore requires the hearer to reconstruct not only a propositional content but an attitude, and to do so on the basis of contextual cues that are frequently implicit rather than lexically signalled.
This attitudinal and context-dependent character is precisely what distinguishes irony from more conventionalised forms of non-literal expression and what makes it a particularly demanding test of pragmatic competence. Where a routinised expression can, in principle, be resolved through stored association, irony typically requires the hearer to detect a mismatch between an utterance and its context, to infer that this mismatch is intentional rather than erroneous, and to recover the evaluative stance that the speaker intends to communicate through that deliberate incongruity. Because this process is inferential and context-bound rather than retrieval-based, irony offers a theoretically privileged window onto the mechanisms of pragmatic comprehension in a way that more conventionalised figurative forms do not.
This same property renders irony comparatively resistant to computational modelling. Large Language Models (L.L.M's) acquire linguistic competence through distributional exposure to text, which allows them to approximate ironic readings that are lexically or contextually well represented in their training data, yet leaves them comparatively exposed when successful interpretation depends on modelling speaker attitude, situational incongruity, or culturally specific expectations that are not straightforwardly encoded in surface form. Because irony requires the reconstruction of speaker intention rather than the retrieval of a stored figurative association, it constitutes a particularly informative test case for evaluating whether L.L.M's approximate human-like pragmatic inference or instead rely on surface-level cues that happen to co-occur with ironic use, such as exaggeration, punctuation, or conventionally sarcastic lexical items.
Moroccan Darija provides a uniquely demanding environment in which to examine this question. As a primarily oral, diglossic variety that coexists with Modern Standard Arabic, French, and Tamazight, Darija makes extensive pragmatic use of irony as a register for social commentary, indirect criticism, and humour, yet this usage is rarely codified in standardised written corpora and is consequently underrepresented in the training data available to most L.L.M's. This combination of high communicative reliance on ironic expression and low textual representation makes Darija a theoretically important test case: a comparison of native-speaker and L.L.M interpretation of irony in this variety speaks not only to Applied Linguistics and Psycholinguistics, but also to Pragmatics, Cognitive Linguistics, and A.I evaluation research concerned with how well language technologies generalise pragmatic inference beyond their dominant training distributions.
2. Research Problem
It is already established that L.L.M's can produce fluent and often contextually plausible responses to ironic and sarcastic language in high-resource languages, and that Darija has been the object of sustained N.L.P development for tasks such as machine translation, sentiment analysis, and dialect identification. What remains unknown is whether this demonstrated fluency reflects an interpretive process comparable to that of native Darija speakers, or whether it reflects the reproduction of statistically frequent cues, such as hyperbole, exclamatory punctuation, or conventionally ironic lexical items, that happen to coincide with correct answers on conventional items while breaking down on items that require genuine contextual and attitudinal inference. This distinction matters academically because it bears directly on competing accounts of L.L.M pragmatic competence: if L.L.M performance tracks human interpretation closely across both conventional and novel irony, this would support claims that emergent language understanding extends to attitudinal, context-driven inference; if performance diverges specifically on items requiring cultural grounding or inferential depth, this would instead support the view that L.L.M competence in low-resource dialectal varieties remains associative rather than inferential. The present study is designed to adjudicate between these two accounts by asking, specifically, whether L.L.M's can interpret ironic meaning in Moroccan Darija in a manner comparable to native speakers.
3. Research Gap
Existing research relevant to this problem falls into two largely separate strands. The first strand consists of N.L.P evaluation studies on irony and sarcasm detection, which have assessed model performance using accuracy- and F.1-based classification metrics, predominantly in high-resource languages such as English and, more recently, in Modern Standard Arabic sentiment contexts; parallel work on Darija has focused on sentiment classification, dialect identification, and machine translation. These studies establish that Darija text can be processed with reasonable technical success but do not examine whether model outputs correspond to a human-like comprehension process, since their evaluation criteria are outcome-based classification metrics rather than process-based measures of interpretive accuracy, agreement, or confidence calibration.
The second strand consists of psycholinguistic studies of irony comprehension, which have investigated the roles of context, common ground, and speaker attitude in English and, to a lesser extent, other well-resourced languages, using controlled experimental paradigms. These studies have not been extended to Darija, and none directly compare human interpretation against L.L.M output using the same controlled, contextually grounded stimuli and coding procedure. No identified study integrates the methodological rigour of psycholinguistic irony-comprehension research with a systematic human–L.L.M comparison in an underrepresented, primarily oral Arabic dialect. This study addresses that intersection directly, offering an empirical basis for evaluating L.L.M pragmatic competence in Darija that goes beyond task-level classification accuracy to item-level, context-sensitive interpretation data.
Restricting the investigation to irony alone, rather than surveying figurative language more broadly, is a deliberate theoretical choice rather than a narrowing of convenience. Irony is distinguished from more conventionalised non-literal forms by its comparatively greater dependence on online, context-driven inference over stored association: its comprehension implicates speaker attitude, common ground, and the detection of situational incongruity in a way that lexically fixed expressions do not. A study that mixed irony with categories resolvable through retrieval would risk conflating two distinct cognitive mechanisms, retrieval and inference, and would obscure which of the two actually drives any observed human–L.L.M divergence. Isolating irony allows the present study to offer a mechanism-specific test of pragmatic inference, yielding theoretical insight that a broader, multi-category design would dilute.
4. Research Aim
The aim of this study is to determine the extent to which Large Language Models' interpretation of ironic utterances in Moroccan Darija corresponds to that of native speakers, and to identify the contextual and linguistic conditions under which this correspondence breaks down. This aim is pursued through the following measurable objectives, each of which maps directly onto a research question and a corresponding methodological procedure:
- Quantify native speakers' interpretation accuracy for conventional and novel ironic utterances, under both full and reduced contextual conditions.
- Quantify L.L.M interpretation accuracy on the identical stimulus set under a standardised prompting protocol applied uniformly across models.
- Compute the level of agreement between human and L.L.M interpretation classifications, using Cohen's kappa as the agreement statistic.
- Statistically compare accuracy and agreement across irony type and contextual explicitness to identify the conditions that produce the greatest human–L.L.M divergence.
- Interpret the resulting accuracy and agreement patterns in light of Relevance Theory, Grice's Cooperative Principle, and the Graded Salience Hypothesis.
5. Research Questions
Main Research Question. To what extent do the interpretations of ironic utterances produced by Large Language Models correspond to those of native Moroccan Darija speakers, and does this correspondence vary systematically with irony type and contextual explicitness?
Sub-questions:
- What proportion of ironic items, differentiated by type (conventional versus novel irony), is correctly interpreted by native speakers compared with each L.L.M?
- What is the level of agreement, measured by Cohen's kappa, between human and L.L.M interpretation classifications, overall and within each irony type?
- Does reducing the contextual information available for an ironic item lower L.L.M interpretation accuracy to a greater extent than it lowers native-speaker accuracy?
- Which contextual cues (for example, explicit evaluative framing versus implicit situational incongruity) are associated with successful irony interpretation for native speakers, and are the same cues associated with successful interpretation for L.L.M's?
- To what extent does self-reported interpretation confidence predict interpretation accuracy for native speakers, and does this relationship hold for L.L.M's?
6. Hypotheses
H.1: Native speakers and L.L.M's will diverge more sharply on culturally dependent, novel irony than on conventional, readily recognisable irony, reflecting the greater reliance of novel irony on situated, context-driven inference rather than stored association; this will be tested using a mixed-effects logistic regression with interpretation accuracy as the outcome, group (human/L.L.M) and irony type as fixed effects, and item and participant as random effects.
H.2: Reducing the contextual information accompanying an ironic item will lower L.L.M interpretation accuracy to a significantly greater extent than it lowers native-speaker accuracy, tested through a group-by-context-condition interaction term in the regression model, reflecting native speakers' greater capacity to reconstruct missing context from world knowledge and shared cultural schemas.
H.3: Native-speaker inter-rater agreement will exceed human–L.L.M agreement, measured by Cohen's kappa, on novel ironic utterances more than on conventional ironic utterances, indicating that model interpretation is comparatively less stable precisely where inferential demands are highest.
H.4: For native speakers, self-reported confidence ratings will positively correlate with interpretation accuracy, reflecting adequate metacognitive calibration; this correlation is predicted to be weaker or absent for L.L.M's, indicating a dissociation between a model's expressed confidence and its actual interpretive accuracy.
H.5: Divergence in interpretation accuracy across the four L.L.M's (gpt, Claude, Gemini, DeepSeek) will be greater for culturally dependent, implicit irony than for explicit, conventional irony, reflecting differential representation of Darija pragmatic norms across each model's training corpus.
7. Theoretical Framework
Each theoretical strand adopted in this study is selected to account for a specific component of the irony construct and analysis, rather than to provide general background. Cooperative Principle and the associated theory of Conversational Implicature supply the foundational account of how irony is recognised in the first instance: an ironic utterance is treated as flouting the maxim of Quality, prompting the hearer to infer an implicature that departs from, and is frequently opposed to, the literal proposition. This framework motivates the study's basic assumption that irony detection is itself an inferential achievement rather than a given, and grounds the binary irony-judgment task included in the Instrument.
Relevance Theory extends and refines this account by treating irony as an echoic use of language, in which the speaker echoes a thought, norm, or expectation while simultaneously distancing herself from it, and the hearer is guided by considerations of cognitive relevance to recover both the echoed content and the speaker's dissociative attitude. This attitudinal dimension is not fully captured by the Gricean account alone and is central to the present study's coding scheme, which distinguishes correct recovery of evaluative polarity from mere detection of non-literalness.
Echoic Theory, developed within the relevance-theoretic tradition and complemented by allusional accounts of irony, is retained as a distinct analytic strand because it specifies the mechanism by which culturally shared expectations become available for echoic use. This is directly relevant to the construction of the Darija stimulus set, since many naturally occurring ironic utterances in the variety depend on echoing community-specific norms, proverbial expectations, or recognisable social scripts that a model without access to Moroccan cultural context would have no basis for reconstructing.
The Graded Salience Hypothesis contributes the processing-level account that motivates the study's distinction between conventional and novel irony. Under this hypothesis, salient meanings, those that are conventional, frequent, or prototypical, are accessed initially regardless of contextual fit, while less salient, novel, or context-dependent meanings require additional inferential effort to compute. This predicts that conventional ironic expressions, being more salient, should be processed comparatively easily by both native speakers and L.L.M's, whereas novel irony, requiring the suppression of a salient literal meaning in favour of a context-derived ironic one, should expose divergence between inference-based human comprehension and association-based model behaviour.
Finally, the Transformer architecture and the distributional semantics it implements provide the framework for interpreting L.L.M behaviour itself, since they predict that model performance should track the frequency and consistency with which a given ironic pattern, or its associated contextual cues, appears in training data, rather than the inferential and attitudinal demands that pattern places on a human comprehender. Read together, these frameworks generate the differentiated, condition-specific predictions tested in the Hypotheses above: a Gricean and relevance-theoretic account of what successful irony comprehension requires, and a distributional account of why L.L.M's might fall short of it under specifiable conditions.
8. Research Design
The study adopts a quantitative, comparative, experimental, cross-sectional design. A quantitative approach is required because the research questions concern the statistical comparison of accuracy and agreement scores across groups and conditions, rather than the qualitative characterisation of individual interpretive processes. The design is comparative because the central research question depends on directly contrasting native speakers with four L.L.M's on an identical stimulus set.
It is experimental in the sense that two theoretically motivated factors are systematically manipulated: irony type (conventional versus novel) and contextual explicitness (full versus reduced context), while surface features such as utterance length and lexical frequency are held approximately constant, allowing both factors to be treated as independent variables rather than confounds. A cross-sectional design is appropriate because the study aims to capture a single comparative snapshot of interpretation performance rather than to track change over time, which is consistent with the aim of establishing whether a correspondence exists at a given point rather than how it develops.
The stimulus set is restricted to irony throughout, categorised along two theoretically justified dimensions rather than mixed with unrelated figurative devices. The first dimension, conventional versus novel irony, is grounded in the Graded Salience Hypothesis and distinguishes routinised ironic expressions with an established figurative status in Darija usage from spontaneously constructed irony requiring online inferential computation. The second dimension, full versus reduced context, is manipulated within item, such that each ironic utterance is presented to different stimulus lists with either a complete contextual scenario or a minimal contextual cue, enabling a direct test of H.2 without confounding context reduction with item identity.
9. Participants
The target population consists of native Moroccan Darija speakers, with a planned sample of 40 to 60 participants. This range is consistent with sample sizes reported in comparable psycholinguistic irony-comprehension studies employing repeated-measures designs, and is expected to provide adequate statistical power to detect medium-sized effects of irony type and contextual condition in the mixed-effects models specified under Hypotheses 1 to 3, given that each participant contributes multiple observations across items. Inclusion criteria require that participants be native speakers of Moroccan Darija, at least 18 years of age, and literate in written Darija, since the task requires reading and written response production. Exclusion criteria include prolonged residence outside Morocco during formative language-acquisition years and self-reported language or reading impairments that could affect comprehension of written stimuli.
Participants will be recruited through a combination of purposive and snowball sampling via university networks and community channels, which is appropriate given the absence of a centralised sampling frame for this population. All participants will provide informed consent prior to participation, and the study protocol will be submitted for approval to the relevant institutional ethics committee before data collection begins.
10. Materials
The stimulus set will consist of approximately 40 to 50 authentic ironic Moroccan Darija utterances, each embedded in a brief contextual scenario establishing the situational incongruity necessary for an ironic reading to arise. Item construction will proceed in two phases. In the generation phase, candidate items will be drawn from naturally occurring Darija discourse, including everyday conversational exchanges and informal social media use, and modelled on these examples by the researcher to ensure ecological plausibility; each item will be assigned to one of two theoretically motivated types, conventional or novel irony, on the basis of its degree of established figurative routinisation in Darija usage. In the validation phase, a separate panel of 10 to 15 native speakers who will not participate in the main study will rate each item for perceived irony, naturalness, type classification, and cultural familiarity on 5-point scales; items will be retained only if they achieve at least 80% panel agreement on their intended type classification and a minimum naturalness rating, and items falling below this threshold will be revised or discarded.
Each retained item will be produced in two contextual versions, a full-context version providing an explicit situational scenario and a reduced-context version providing only a minimal situational cue, with items distributed across two counterbalanced stimulus lists so that no participant, and no L.L.M session, encounters both versions of the same item. Cultural familiarity, obtained from the validation panel, will be retained as an item-level covariate in the statistical models, allowing the study to examine whether human–L.L.M divergence is attributable to cultural specificity independently of irony type. A small set of non-ironic filler items, drawn from ordinary sincere Darija utterances, will additionally be embedded among the ironic stimuli for the sole purpose of supporting the validity of the irony-detection task described in the Instrument section below, by preventing a default response bias toward classifying every item as ironic; these fillers are not treated as a distinct figurative category and are excluded from the main comparative analyses. As an illustration of the kind of item the validated set will contain, an utterance such as [Arabic text]{dir="rtl"} (“the weather is lovely today”), produced in a contextual scenario describing a sudden downpour, exemplifies the echoic, dissociative structure that the coding scheme is designed to capture; such examples are offered here for illustration only and are not themselves assumed to meet the validation threshold described above.
11. Instrument
The instrument is designed to yield four operationally defined dependent variables: irony-detection judgment, interpretation accuracy, confidence rating, and agreement score. The study will be administered through an online survey platform, with each item presented individually in randomised order. For every item, participants will first indicate whether the utterance is ironic or sincere; for items judged ironic, participants will then state, in their own words, the meaning and attitude they believe the utterance conveys, select the closest paraphrase from a small set of plausible alternatives, and rate their confidence in that interpretation using the 5-point scale shown below.
Table summary: A five point confidence scale for perceiving irony. A score of 1 indicates the person is not at all confident, meaning irony was not perceived and no interpretation is available. Scores increase in confidence from 2, slightly confident, to 3, moderately confident, and 4, confident. The highest score of 5 represents being completely confident, where the interpretation was reached without hesitation.
Open-ended responses will be coded into five mutually exclusive categories using the rubric below, which operationalises interpretation accuracy for both human and L.L.M data.
: Table summary: Operational definitions for categorizing ironic interpretations. A correct ironic interpretation recovers the speaker-intended attitudinal meaning and evaluative polarity, such as mocking or critical. A partially correct interpretation captures the general polarity or dissociative stance but misses specific nuances or targets. Incorrect interpretations are unrelated to or inconsistent with the intended reading, while literal misinterpretations treat the utterance as a sincere assertion. Finally, culturally inappropriate interpretations provide a plausible non-literal reading that is inconsistent with Moroccan usage norms.
Coding will be performed independently by two trained raters who are blind to whether a given response was produced by a human participant or by an L.L.M, and inter-rater reliability will be established using Cohen's kappa, with a minimum threshold of κ = 0.75 required for acceptable agreement; disagreements will be resolved through discussion or, where necessary, adjudication by a third rater. The agreement score between human and L.L.M responses, the study's primary comparative dependent variable, will itself be computed as a second application of Cohen's kappa, comparing the coded response category assigned to each item's human response with that assigned to the corresponding L.L.M response. As an optional secondary variable, response latency will be recorded via the survey platform's timestamp function and analysed as a proxy for processing effort, consistent with accounts predicting that novel, reduced-context irony should elicit longer response times than conventional, fully contextualised irony.
12. Human Data Collection
Human data collection will proceed in four stages over an estimated two-week period. First, a pilot test will be conducted with 8 to 10 participants drawn from the target population but excluded from the main sample, in order to verify that instructions are unambiguous, that the irony-detection and interpretation tasks are clearly understood, and that the response format yields codeable data; the instrument will be revised in response to pilot feedback before full deployment. Second, recruited participants will complete an informed consent form outlining the purpose of the study, the voluntary nature of participation, and data-handling procedures, followed by a short demographic questionnaire capturing age, region, and language background.
Third, participants will be randomly assigned to one of the two counterbalanced stimulus lists and will complete the main interpretation task, preceded by a worked example item to standardise task understanding across participants. Fourth, collected responses will be screened for completeness, straight-lining, and evidence of non-engagement, with flagged cases excluded prior to analysis. The full protocol, including consent procedures and data storage arrangements, will be submitted for institutional ethics approval prior to the pilot stage.
13. L.L.M Data Collection
L.L.M data will be collected from four models, gpt, Claude, Gemini, and DeepSeek, each accessed via its official A.P.I to ensure a stable and documentable configuration; the exact model version for each system, together with the date of data collection, will be recorded in the final report, since model behaviour can change across releases. The temperature parameter will be set to 0 across all four models to maximise output determinism and reduce uncontrolled response variability, and all other generation parameters will be left at their documented default values and reported in full for each model. Each stimulus item will be submitted to each model in a separate, independent A.P.I session to prevent carry-over effects from preceding items, and the prompt wording will be standardised and held identical across all four models to closely parallel the instructions given to human participants, including an equivalent worked example, to preserve comparability across both models and conditions.
To assess the stability of L.L.M output, a subset of items will be resubmitted three times in independent sessions for each model; instability in the coded response category across repetitions will be reported descriptively and treated as a limitation of single-response comparisons. All raw L.L.M outputs will be recorded verbatim, together with the exact prompt, model identifier, and timestamp for each query, and will then be coded using the identical rubric and blind procedure applied to human responses, ensuring that all datasets, human and model, are directly comparable at the level of both measurement and analysis.
14. Statistical Analysis
The primary analyses will compare interpretation accuracy between native speakers and each L.L.M using mixed-effects logistic regression, with group, irony type, and context condition entered as fixed effects and their interactions specified to test Hypotheses 1 and 2, and item and participant entered as random effects to account for repeated measurement. Human–L.L.M agreement will be quantified using Cohen's kappa, computed overall and separately within irony type and context condition, and compared descriptively across the four models to address H.3 and H.5; inter-model agreement will additionally be examined using Fleiss' kappa across the full four-model set. The relationship between self-reported confidence and interpretation accuracy will be modelled separately for native speakers and for each L.L.M using logistic regression, with the resulting coefficients compared to test the metacognitive-calibration prediction in H.4. Cultural familiarity, obtained during stimulus validation, will be entered as an item-level covariate in all models to examine whether it accounts for variance in human–L.L.M divergence independently of irony type.
Where multiple comparisons are conducted across the four L.L.M's, a false discovery rate correction will be applied to control the familywise error rate. All analyses will be conducted with an alpha level of 0.05, and effect sizes will be reported alongside significance tests to support interpretation of practical, and not merely statistical, significance.
15. Expected Contributions
This study is expected to contribute to four interrelated areas. Theoretically, it offers an empirical test of Gricean and relevance-theoretic accounts of irony comprehension against a distributional, association-based account of L.L.M behaviour, using naturalistic, contextually grounded dialectal materials rather than constructed or translated stimuli. For pragmatic theory more broadly, the explicit manipulation of contextual explicitness alongside irony type will yield fine-grained evidence on the conditions under which context reduction disproportionately impairs inference-based, as opposed to association-based, comprehension. For research on Moroccan Darija, the study will produce a validated dataset of contextually embedded ironic utterances, classified by type and rated for naturalness and cultural familiarity, constituting a reusable resource for future work in dialectal pragmatics and Arabic N.L.P. Finally, for A.I evaluation research, the study offers a process-level, rather than purely outcome-based, benchmark for assessing whether current frontier L.L.M's perform genuine pragmatic inference or associative pattern matching when interpreting irony in an underrepresented, primarily oral language variety, with direct relevance to the responsible evaluation and deployment of these systems in low-resource linguistic communities.
You have reached the end of the document.