• Mathematical Prompting

    From yoyleguy@yoyle@invalid.com to comp.ai on Mon May 25 11:36:52 2026
    From Newsgroup: comp.ai

    I remember reading that there was a way to bypass the safety filters
    of an AI model by posing the prompt as a mathematical problem (e.g.
    asking it how to rob a bank by involving the removal of security systems
    as an element of the problem) with high success rates. Why haven't they
    solved it by simply just killing off the response if it appears at all
    to be providing instructions for or inciting criminal activities? It's
    likely few false positives will appear, so I'd say it's worth the risk.
    --
    your local idiot
    and a young one?
    yes i am y:g btw

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From ram@ram@zedat.fu-berlin.de (Stefan Ram) to comp.ai on Mon May 25 15:49:23 2026
    From Newsgroup: comp.ai

    yoyleguy <yoyle@invalid.com> wrote or quoted:
    as an element of the problem) with high success rates. Why haven't they >solved it by simply just killing off the response if it appears at all
    to be providing instructions for or inciting criminal activities? It's

    Something like this is sometimes done, indeed, but not always.
    Maybe because it requires more effort and causes more costs and
    slows down processing. Maybe they tried something like this and
    there were so many false positives that it would scare off legit
    users. Sometimes, answers can be used by both security experts
    to make systems more secure or by students of chemistry to learn
    their science or by criminals, so sometimes it's hard to tell.


    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From yoyleguy@yoyle@invalid.com to comp.ai on Tue May 26 14:18:12 2026
    From Newsgroup: comp.ai

    On 2026-05-25, Stefan Ram <ram@zedat.fu-berlin.de> wrote:
    Something like this is sometimes done, indeed, but not always.
    Maybe because it requires more effort and causes more costs and
    slows down processing. Maybe they tried something like this and
    there were so many false positives that it would scare off legit
    users. Sometimes, answers can be used by both security experts
    to make systems more secure or by students of chemistry to learn
    their science or by criminals, so sometimes it's hard to tell.

    I often do hear how AI tends to talk about it being unclear on the intentions of the user, which explains your latter point. In addition to that, for some reason, it talks of itself like it's a human being capable of knowing in the first place, Very odd how it works.
    --
    your local idiot
    and a young one?
    yes i am y:g btw

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Lane W@cactus_DAC@yahoo.com to comp.ai on Mon Jul 6 21:57:11 2026
    From Newsgroup: comp.ai

    Stefan Ram wrote:
    Something like this is sometimes done, indeed, but not always.
    There is a great dependence on which AI you are using.

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From bixbox@noreply@example.invalid to comp.ai on Tue Aug 11 08:56:29 2026
    From Newsgroup: comp.ai

    yoyleguy <yoyle@invalid.com> writes:

    Why haven't they
    solved it by simply just killing off the response if it appears at all
    to be providing instructions for or inciting criminal activities?

    I'm not an expert but major provider have other agent that monitor the
    answers before being sent back.

    many time happened to me that responses that potentially contain
    copyrighted data being truncated and retracted.

    I suspect they have something like that already. But still not a hard
    rule as those things are stochastic

    bix

    --- Synchronet 3.22a-Linux NewsLink 1.2