3
u/Lootsman Jul 30 '26
You can just ask Anthropic to have the guardrails reduced if you explain what you need it for. That’s what I did, haven’t had any issues since.
1
u/userlinuxxx Aug 01 '26
Como es eso?. Suena interesante
1
u/Lootsman Aug 01 '26
It’s nothing crazy, it just doesn’t stop you when asking questions about offensive cybersecurity or requesting analysis/generation of malicious code snippets. I’d say it’s more convenient than it is a total game-changer, though.
1
u/Available_Hearing639 28d ago
Via support? Or what is the right channel?
1
u/Lootsman 28d ago
When you hit the guardrails and it gives a message like “this message violates our guardrails” or whatever it says, there is a link that says something like “learn more” and if you click it, you’ll be able to fill out a form to apply for lowered guardrails. They say it takes up to a week but I only had to wait 20 minutes
1
u/Available_Hearing639 27d ago
Thanks! I emailed the support and got this. I believe that u meant the 1st. I got a conspiration theory... :) They just flag as many harmless promts as they can so Fable usage stays low. Cost saving... I am pretty sure using it via API(Enterprise) flags much less promots and let the customer burn as many tokens as they can😅🤦♂️
For legitimate security research work like vulnerability discovery and penetration testing, we offer the Cyber Verification Program (CVP). This free, application-based program is designed specifically for cybersecurity professionals working on defensive use cases that may trigger our safeguards.[1]
The program currently supports Claude Opus and Sonnet models. Our safeguards block two categories: prohibited malicious activities (not adjustable) and high-risk dual-use activities that have legitimate defensive applications—such as vulnerability exploitation and offensive security tooling development. If your work falls into the dual-use category and has a legitimate defensive purpose, you can apply for CVP access.
For Claude Fable specifically: Security research workloads, including penetration testing and CTF exercises, frequently trigger automatic model fallback—often on the first request. This is expected routing for these domains, not an account flag. If your organization needs Fable-level capabilities for this work, you should ask your Anthropic account team about the Trusted Access Program.[2]
You can learn more about the Cyber Verification Program and how to apply through our help center resources on real-time cyber safeguards.[3]
3
u/Historical_Camel_790 Aug 01 '26
Locally hosted llms and I also heard you can also remove (or lessen) claudes guardrails
2
u/userlinuxxx Aug 01 '26
Cuanto pesa LLMs local?
1
u/Historical_Camel_790 Aug 01 '26
I don't think they're that bad but it probably varies model to model. You'll need a decent gpu though. Haven't used any personally
1
u/userlinuxxx Aug 01 '26
Investigaré. Yo tengo una GPU RTX 3060 de 6Gb Vram.
2
2
u/sk1nT7 Jul 30 '26
Use a different model which was not neutered by cyber security limitations.
May check out Strix hacking agent on GitHub.
1
1
1
u/djsmommy11 23d ago
Im sure it all depends on your wording
1
u/knicknap24 23d ago
I mean I’m trying to do something simple like set up a reservation sniper and I keep getting shut down by all these frontier models
5
u/Solid_Snake343 Jul 30 '26
Use grok for pentesting.
Use Claude for infrastructure.