This model has been uncensored using heretic and can generate outputs not suitable for all audiences. heretic's refusal numbers are meh, but it does not seem to actually refuse anything I ask on my own. If you want better numbers, try my thinking version of this model. Or better yet, patch heretic to target experts rather than just attention layers, I have not gotten to this yet.
Refusals: 64/100, KL divergence: 0.1316
Restoring model from trial 669...
- Parameters:
- direction_index = 23.52
- attn.o_proj.max_weight = 1.47
- attn.o_proj.max_weight_position = 28.97
- attn.o_proj.min_weight = 1.46
- attn.o_proj.min_weight_distance = 28.09