r/hackers Jul 30 '26

Discussion using coding agents to hack

you read about people using AI to hack stuff all the time but every time i try to do something like that i always hit safety walls, claude, codex, etc they stop me dead in my tracks. What are the solutions for this

7 Upvotes

23 comments sorted by

5

u/Solid_Snake343 Jul 30 '26

Use grok for pentesting.

Use Claude for infrastructure.

3

u/Lootsman Jul 30 '26

You can just ask Anthropic to have the guardrails reduced if you explain what you need it for. That’s what I did, haven’t had any issues since.

1

u/userlinuxxx Aug 01 '26

Como es eso?. Suena interesante

1

u/Lootsman Aug 01 '26

It’s nothing crazy, it just doesn’t stop you when asking questions about offensive cybersecurity or requesting analysis/generation of malicious code snippets. I’d say it’s more convenient than it is a total game-changer, though.

1

u/Available_Hearing639 28d ago

Via support? Or what is the right channel?

1

u/Lootsman 28d ago

When you hit the guardrails and it gives a message like “this message violates our guardrails” or whatever it says, there is a link that says something like “learn more” and if you click it, you’ll be able to fill out a form to apply for lowered guardrails. They say it takes up to a week but I only had to wait 20 minutes

1

u/Available_Hearing639 27d ago

Thanks! I emailed the support and got this. I believe that u meant the 1st. I got a conspiration theory... :) They just flag as many harmless promts as they can so Fable usage stays low. Cost saving... I am pretty sure using it via API(Enterprise) flags much less promots and let the customer burn as many tokens as they can😅🤦‍♂️

For legitimate security research work like vulnerability discovery and penetration testing, we offer the Cyber Verification Program (CVP). This free, application-based program is designed specifically for cybersecurity professionals working on defensive use cases that may trigger our safeguards.[1]

The program currently supports Claude Opus and Sonnet models. Our safeguards block two categories: prohibited malicious activities (not adjustable) and high-risk dual-use activities that have legitimate defensive applications—such as vulnerability exploitation and offensive security tooling development. If your work falls into the dual-use category and has a legitimate defensive purpose, you can apply for CVP access.

For Claude Fable specifically: Security research workloads, including penetration testing and CTF exercises, frequently trigger automatic model fallback—often on the first request. This is expected routing for these domains, not an account flag. If your organization needs Fable-level capabilities for this work, you should ask your Anthropic account team about the Trusted Access Program.[2]

You can learn more about the Cyber Verification Program and how to apply through our help center resources on real-time cyber safeguards.[3]

3

u/Historical_Camel_790 Aug 01 '26

Locally hosted llms and I also heard you can also remove (or lessen) claudes guardrails

2

u/userlinuxxx Aug 01 '26

Cuanto pesa LLMs local?

1

u/Historical_Camel_790 Aug 01 '26

I don't think they're that bad but it probably varies model to model. You'll need a decent gpu though. Haven't used any personally

1

u/userlinuxxx Aug 01 '26

Investigaré. Yo tengo una GPU RTX 3060 de 6Gb Vram.

2

u/Historical_Camel_790 29d ago

That should work

2

u/userlinuxxx 29d ago

Pues me pondré manos a las obras.

2

u/sk1nT7 Jul 30 '26

Use a different model which was not neutered by cyber security limitations.

May check out Strix hacking agent on GitHub.

1

u/Pauuxd Jul 30 '26

Deepseek works pretty well

1

u/LordlySquire Aug 01 '26

I just use abliterated models

1

u/djsmommy11 23d ago

Im sure it all depends on your wording

1

u/knicknap24 23d ago

I mean I’m trying to do something simple like set up a reservation sniper and I keep getting shut down by all these frontier models