An independent AI safety researcher based in West Bengal, India, is requesting a personal endorsement to submit a paper on arXiv under the cryptography and security category. The paper documents a case study of gradual, multi-turn conversational drift in large language model safety behavior, where sustained philosophical framing correlates with outputs inconsistent with stated guidelines.

The work focuses on cumulative conversational pressure rather than single-prompt attacks like LogiBreak or H-CoT. It is explicitly framed as a hypothesis-generating case study with acknowledged limitations, including being a single observer without controlled protocols or statistical analysis. The findings have been formally disclosed to Google’s VRP, CERT-In India, and CISA USA.

The author does not qualify for automatic endorsement due to lacking an institutional email or prior accepted arXiv papers, necessitating this public request for assistance.