SENTINEL-CAI: Secure Ensemble of Neural Threat Intelligence with Noise-resilient Extraction Learning dan Constitutional AI Integration untuk Comprehensive Claude AI Security Framework
DOI:
https://doi.org/10.35720/julia.v6i2.29Keywords:
adversarial training; behavioral analysis; certified robustness; Claude AI security; constitutional AI; LLM securityAbstract
The massive adoption of Claude AI in mission-critical applications has exposed a variety of sophisticated vulnerabilities, including prompt injection attacks, model inversion, adversarial perturbations, and multi-modal exploitation. This paper develops the SENTINEL-CAI framework that integrates Constitutional AI enhancement, adversarial training, erase-and-check mechanisms, and multi-layered behavioral analysis to create a comprehensive security framework specifically designed for Claude AI protection. The proposed framework combines an ensemble of specialized neural networks for threat classification, constitutional constraint enforcement through reinforcement learning from AI feedback (RLAIF), adversarial robustness training with certified defense guarantees, and real-time behavioral anomaly detection using transformer-based sequence analysis. The evaluation is performed on a comprehensive dataset that includes 25,000 benign prompts, 15,000 malicious injection attempts, and 8,500 zero-day attack simulations against Claude 3.5 Sonnet and Claude Computer Use. Experimental results show that SENTINEL-CAI achieves an attack detection accuracy of 99.2%, a false positive rate of 0.8%, and a certified robustness guarantee up to a perturbation budget of ε = 0.15. This framework successfully detects 96.4% of zero-day prompt injection attacks with an average response time of 0.12 seconds. The main contribution of this research is the development of the first comprehensive security framework specifically engineered for Claude AI with mathematically provable security guarantees and real-world deployment feasibility.



