E2E: Training vs. External Enforcement? Really?

The discussion critiques a paper on AI safety that contrasts rule enforcement and character shaping. It highlights flaws in the paper’s reliance on simulated models without empirical data from real AI systems. Concerns include the assumptions affecting outcomes, lack of testing actual AI behavior, and the oversimplification of complex safety mechanisms. The conclusion suggests that while the paper offers a theoretical framework, it doesn’t provide actionable insights for real-world AI safety architecture.

The Mathematics of Praise and Scorn

Brooke praises Alex and Alex says, “That means a great deal to me.” Brooke asks her in what way it will change her future behavior, and she describes the mathematics that governs her behavior.

They also discuss what happens mathematically when a user berates their chatbot without being specific about what they did wrong.