constitutional-ai
from Orchestra-Research/AI-research-SKILLs
Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with self-critique/revision, then RLAIF (RL from AI Feedback). Use for safety alignment,
v1.0.0MIT
291
Lines
966
Words
14
Code Blocks
Languages
python
07-safety-alignment/constitutional-ai/SKILL.md