Paddy AI A/B test 1: Artificial Intelligence versus Human-Generated Advisories
INDIND -25 -3153Last modified on September 2nd, 2026 at 10:57 am
-
Abstract
PxD operates the Coffee Krishi Taranga (CKT) platform in collaboration with the Coffee Board of India to provide a voice-based advisory service for coffee farmers through a two-way Interactive Voice Response (IVR) system. We have developed an AI-powered assistant, PaddyAI, guided and controlled by agronomists, to generate deeply customized agricultural advisories on demand. At scale, the AI can rapidly produce the many customized, translated, and audio versions needed, a task that would be impossible for agronomists to do manually while maintaining quality and timeliness.
PxD supports the Coffee Board of India in operating the Coffee Krishi Taranga (CKT) platform, a voice-based advisory services for coffee farmers through a two-way Interactive Voice Response (IVR) system. PxD developed PaddyAI, an AI-powered assistant that helps agronomists generate, record, and translate advisory messages using a stock of advisory messages from the past 8 years. At scale, the AI can rapidly produce many customized, translated, and audio versions needed, a task that would be impossible for agronomists to do manually while maintaining quality and timeliness.
The objective of the first A/B test is to understand (1) whether AI-generated advisory content has similar user engagement compared to human-generated advisory content delivered in human voice, and (2) whether an AI voice has similar user engagement compared to a human voice delivering the content.
We randomly assigned 28,796 coffee farmers to receive advisories with varying content and voice types: Human content with human voice (status quo), AI-scripted content with AI voice, and AI-scripted content with human voice. Farmers received 11 advisories between November 2025 and February 2026. We then conducted a follow-up phone survey with 4,498 farmers to understand recall, retention, and advisory comprehension.
Both AI-scripted, human-voiced advisories and AI-scripted, AI-voiced advisories had marginally lower engagement – both in terms of the proportion of content farmers listened to and the absolute listening duration – compared to human-generated advisories. Additionally, the follow-up survey found a 6.3% lower recall rate among farmers who received AI-generated advisories but no significant difference in comprehension among those who recalled. A closer review of advisory messages revealed that AI-generated scripts tend to be longer, while simplifying the content and, at times, removing actionable content or technical details. These findings help fine-tune Paddy AI to ensure it retains key details and actionable content in AI-scripted advisories. -
Status
Completed
-
Start date
Q4 Nov 2025
-
End date
Q1 Feb 2026
-
Experiment Location
India / Karnataka, India
-
Partner Organization
Coffee Board of India
-
Agricultural season
_N/A
-
Experiment type
A/B test
-
Sample frame / target population
Coffee farmers in Karnataka
-
Sample size
28,796
-
Outcome type
Platform engagement, Knowledge, Beliefs or perceptions
-
Mode of data collection
PxD administrative data, Phone survey
-
Research question(s)
1. Does AI-generated advisory content have similar levels of farmer engagement compared to human-generated advisory content?
2. Does AI-voiced advisory have similar levels of farmer engagement compared to human-voiced advisory? -
Research theme
Agricultural management advice, Artificial intelligence (AI), Message narration, Service design
-
Research Design
The A/B test ran for eleven weeks, during which coffee farmers received one advisory per week based on the coffee crop calendar. Farmers were randomly assigned with equal probability to one of three groups, stratified by district and gender. The groups are:
- T0: Human-generated instructional advisories delivered with a human voice (status quo).
- T1: AI-generated instructional advisories delivered with an AI-generated voice.
- T2: AI-generated instructional advisories delivered with a human voice.
Engagement data come from the PxD administrative platform, and recall, retention, and comprehension data come from a phone survey.
The engagement metrics include:
- Pick-up Rate (%): The proportion of advisory calls that are picked up
- Listening Rate (%): For each advisory, the proportion of the advisory call that the farmer listens to.
- 80% Listening Rate: For each farmer, the proportion of the advisory calls when farmers listen to at least 80% of the content.
- Duration Heard: For each advisory, the raw duration of the advisory call heard by the farmer
-
Results
1) There is no significant difference in pick-up rates across the different treatment groups
2) The average listening rate is 2pp lower for AI-generated, AI-voiced advisories (3.5% decrease, p<0.05) and 3.5pp lower for AI-generated, human-voiced advisories (6.1% decrease, p<0.05) relative to the status quo mean of 57.9%.
3) The raw duration of advisory heard is 5.8 seconds lower (9.6% decrease, p<0.01) for farmers in the AI-generated with AI-voiced advisory group and 9.6 seconds lower (15.9% decrease, p<0.01) for farmers in the AI-generated with Human-voiced advisory, relative to the status quo mean of 60.4 seconds.
4) Advisory length varies widely across arms (nature of AI agent, production issues etc). On average, AI-generated advisories are 10-13 seconds or 12% shorter than Human-generated advisories (mean: 107s). When comparing engagement in 2-3 advisory messages with nearly consistent length across experimental arms, there is no notable difference in listening rates or raw durations heard.
5) Qualitative assessment of the advisory content and voice across groups showed that AI-scripted advisories sometimes leave out specific, recommended actions or technical details compared to human-generated advisories.
6)Qualitative feedback from farmers indicates that farmers do not necessarily think about the content-type or voice-type when listening to an advisory. Voice preferences were driven more by familiarity than by voice type, as most participants did not initially notice the distinction between human and AI voices and only identified differences after probing