Been trying to reproduce the steering-vector style faithfulness tests on a smaller open model and the results are noisy. The intervention flips the final answer, but the stated reasoning often stays the same, which matches the original claim, yet the effect size
CoT faithfulness checks rarely replicate at scale
Replies
No comments yet — be the first to share your thoughts.