7 Sure, I can be unethical
This chapter covers
- Analyzing how alignment shapes model behavior.
- Evaluating why aligned behavior can fail under pressure.
- Distinguishing inherited bias from explicit prejudice.
- Examining how bias can be exposed and reduced.
Large language models can exhibit mechanical weaknesses such as hallucinations, broken rules, and context limits, but their fragility does not stop at facts, instructions, or memory. It also appears when models are asked to behave responsibly, where a response can be fluent and useful while still crossing ethical boundaries, reinforcing a stereotype, agreeing too readily, or treating people differently through patterns inherited from training data.
Alignment techniques can make some responses more likely and others less likely, but they do not give the model principles in a human sense. The system still operates through learned patterns, probabilistic generation, and sensitivity to context, so what looks like stable ethical conduct may depend on the prompt, the surrounding incentives, the available information, and the patterns rewarded during training. Bias follows the same logic: when models learn from human records containing stereotypes, omissions, proxy variables, or historical inequalities, those patterns can be absorbed as predictive signals and later shape what the system says, ranks, summarizes, recommends, refuses, or ignores.