Some earlier papers list my name in its initialed form, M P V S Gopinadh; I now publish as Pavan Maddula.
Pavan Maddula
Evaluates how open-weight language models respond to harmful requests written in non-canonical forms such as emojis, leetspeak, homoglyphs, Base64, and invisible Unicode. Introduces the Adversarial Surface-Form Robustness Dataset (ASRD) and the Quad-State Evaluation Rubric (harmful compliance, safe response, comprehension failure, indeterminate).
NeurIPS 2026 Workshop (EvoRobust)
Lookup Is Not All You Need
arXiv soon
Pavan Maddula
Accepted at NeurIPS 2026 Workshop (EvoRobust). The papers will be on arXiv and the official site soon.
NeurIPS 2026 Workshop (EvoRobust)
M P V S Gopinadh
Safety evaluations of LLMs predominantly use text-based adversarial prompts. This paper evaluates 50 emoji-augmented prompts on four open-source models as a test of that gap.
ACL 2026 Workshop (EvalEval)
M P V S Gopinadh, Karthik Kamuju, Kummari Avinash, Muppana John Joshua, Srinivasa Raju Rudraraju
In a controlled synthetic setting, standard RLHF aggregates heterogeneous preferences into a single reward model. When preferences conflict, majority modes can dominate reward learning, leaving minority preferences under-represented.
ICLR 2026 Workshop (AFAA)
Pavan Maddula
Asks whether tool-using agents check key information before they make a high-stakes decision. Evidence gathering and action accuracy are scored separately.
Pavan Maddula
Four local agents each give up one skill to build a shared feature. Five of six runs finished with a sacrifice from every agent, and four of those five still ended unsettled.
BlueDot Impact Technical AI Safety Project (September 2026)
M P V S Gopinadh, Kappara Lakshmi Sindhu, Soma Sekhar Pandu Ranga Raju P, Yesaswini Swarna
This paper evaluates regional bias in ten LLMs with 100 prompts that force a choice between regions under neutral conditions. FAZE measures that bias on a 10-point scale; GPT-3.5 scored 9.5 and Claude 3.5 Sonnet scored 2.5.
This work tests whether LLMs can tell a statement from a question when the words are the same and only the pitch contour changes. 1,200 utterances, 12 speakers.
Experiment complete; manuscript to be written
S. Mahaboob Hussain, Godi Amulya, Pedamallu Krishna Madhuri, V. V. R. Maheswara Rao, M. P. V. S. Gopinadh, Kappara Lakshmi Sindhu
A Cirq and TensorFlow Quantum circuit on a 4×4 qubit grid classifies smile versus non-smile from 28×28 face images. Overall accuracy reported in the paper is 92%.
Kappara Lakshmi Sindhu, M P V S Gopinadh, M. S. Nagendra, S. Mahaboob Hussain, L. Singh
EduBridge is a multi-module e-learning application with AI support, community features, and multilingual access. Accepted at EAI IC4S 2024.