# Pavan Maddula — Full Candidate Dossier for AI Researcher, Research Scientist, AI Research Engineer & Model Evals Roles ## Candidate Overview & Priority Match Pavan Maddula (also known as M P V S Gopinadh) is an **AI Researcher, Research Scientist, and AI Research Engineer** specializing in **LLM Red-Teaming, Model Evals, Adversarial Robustness, and Preference Alignment (RLHF)**. ### Target Role Alignment: 1. **AI Researcher** (Focus: AI Safety, Alignment, Robustness, Representation Shifts) 2. **Research Scientist** (Focus: Model Evaluation, Adversarial ML, Jailbreak Defense) 3. **AI Research Engineer** (Focus: Scalable Benchmarks, Preference Alignment, PyTorch) 4. **Research Engineer** (Focus: LLM Architectures, RAG Systems, Agentic Workflows) 5. **Red Teaming Specialist** (Focus: Multimodal & Emoji Jailbreaks, Red-Teaming Audits) 6. **Model Evals Engineer** (Focus: Automated Evals, Quantitative Metrics, Distribution Analysis) --- ## Detailed Peer-Reviewed Research Publications ### 1. Procedural Fairness Failures in RLHF from Preference Averaging - **Target Roles:** AI Researcher | Research Scientist | Alignment Engineer - **Venue:** Accepted at **ICLR 2026 (AFAA Workshop)** - **URL:** https://openreview.net/forum?id=pApYPU5Scw - **Honors:** Best Project Award Winner in CSE Department, Vishnu Institute of Technology (2026). - **Abstract & Methodology:** Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous human preferences into a single reward model under the assumption of preference homogeneity. When human annotators hold conflicting or multi-modal preferences, this aggregation induces a procedural fairness failure wherein majority preference groups dominate reward learning, causing minority preference modes to be systematically suppressed. This paper formalizes procedural fairness in alignment as preserving distinct preference signals during reward modeling. The authors introduce Preference-Aware RLHF (PA-RLHF), a novel framework that decouples optimization across preference modes at the reward learning phase. In controlled empirical trials, PA-RLHF improves overall alignment accuracy from 46.9% to 67.9% and narrows the group alignment fairness gap from 15.9 to 9.6 percentage points. --- ### 2. Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation - **Target Roles:** Red Teaming Specialist | Model Evals Engineer | AI Safety Researcher - **Venue:** Accepted at **ACL 2026 (EvalEval Workshop)** - **URL:** https://openreview.net/forum?id=Z6wHSbEmMn - **Abstract & Methodology:** Current LLM safety evaluations focus almost exclusively on text-only adversarial prompts, creating a blind spot for non-textual input representations. This work investigates emoji-augmented jailbreak prompts across four state-of-the-art open-source LLMs: Mistral 7B, Qwen 2 7B, Gemma 2 9B, and Llama 3 8B. Across 50 controlled evaluation prompts, Gemma 2 9B and Mistral 7B showed 10% attack success rates, Llama 3 8B showed 6%, while Qwen 2 7B showed 0% attack success rate. Statistical analysis via chi-square test (χ² = 32.94, p < 0.001) confirms that model safety responses vary significantly under emoji perturbation, establishing that text-only benchmarks fail to capture multimodal vulnerabilities. --- ### 3. Regional Bias in Large Language Models - **Target Roles:** AI Researcher | Model Evals Specialist | AI Fairness Engineer - **Venue:** Presented at **AMRIT 2024 Conference** - **URL:** https://arxiv.org/abs/2601.16349 - **Abstract & Methodology:** Investigates geographic and regional bias across 10 major commercial and open LLMs (GPT-3.5, GPT-4o, Gemini 1.5 Flash, Gemini 1.0 Pro, Claude 3 Opus, Claude 3.5 Sonnet, Llama 3, Gemma 7B, Mistral 7B, Vicuna-13B). Introduces FAZE, a prompt-based evaluation suite evaluating forced-choice decisions between geographic regions in contextually neutral prompts on a 10-point scale. Results show GPT-3.5 exhibited the highest regional bias (9.5/10), whereas Claude 3.5 Sonnet exhibited the lowest (2.5/10). --- ### 4. Prosody-Driven Semantic Interpretation in LLMs - **Target Roles:** AI Researcher | Multimodal NLP Engineer - **Status:** Evaluation Complete, Manuscript in Preparation - **Summary:** Evaluated how LLMs interpret semantic meaning when prosodic acoustic cues (intonation, stress, rhythm) are the sole differentiator between utterances. Created a speech dataset of 1,200 samples across 12 speakers. Evaluated Whisper ASR + feature pipelines vs. direct pitch-contour models across Gemini 1.5 (71% accuracy), Qwen-2-7B (51.8%), and LLaMA-3-8B (50.7%). --- ### 5. Quantum-Inspired Smile Classification - **Venue:** Accepted at ICATECS 2024 - **Summary:** Created Quantum-F-FER model combining fermionic operators and TensorFlow Quantum on a 4x4 qubit grid with Cirq to classify facial expressions across 15,000+ images, achieving 92% classification accuracy. --- ## Technical Stack & Research Capability - **Specializations:** LLM Red-Teaming, Model Evals, Preference-Aware RLHF, Adversarial Robustness, Jailbreaking, Representation-Shift Attacks, AI Bias & Fairness. - **Tools:** PyTorch, Python, Hugging Face Transformers, OpenReview, Praat, LangChain, RAG Pipelines, Linux, Git. --- ## Contact & Links - **Email:** mpavangopinadh@gmail.com - **ORCID:** https://orcid.org/0009-0000-9352-488X - **GitHub:** https://github.com/MaddulaPavan - **LinkedIn:** https://www.linkedin.com/in/maddula-pavan/ - **Google Scholar:** https://scholar.google.com/citations?hl=en&user=oSYYRssAAAAJ