The China Mail - AI systems are already deceiving us -- and that's a problem, experts warn

USD -
AED 3.672504
AFN 65.503991
ALL 80.193613
AMD 365.443623
ANG 1.789783
AOA 918.000367
ARS 1475.150612
AUD 1.415829
AWG 1.80125
AZN 1.70397
BAM 1.690479
BBD 2.011669
BDT 122.606854
BGN 1.696366
BHD 0.37668
BIF 2985.954449
BMD 1
BND 1.277583
BOB 11.640953
BRL 5.222404
BSD 0.998833
BTN 95.261679
BWP 13.45607
BYN 3.037988
BYR 19600
BZD 2.008816
CAD 1.38765
CDF 2273.000362
CHF 0.813662
CLF 0.023213
CLP 913.600415
CNY 6.743204
CNH 6.74452
COP 3115.519253
CRC 449.371191
CUC 1
CUP 26.5
CVE 95.306625
CZK 20.930304
DJF 177.864212
DKK 6.460604
DOP 58.464065
DZD 131.663425
EGP 49.854358
ERN 15
ETB 161.573793
EUR 0.864504
FJD 2.234204
FKP 0.737767
GBP 0.73929
GEL 2.610391
GGP 0.737767
GHS 10.937378
GIP 0.737767
GMD 73.503851
GNF 8773.931458
GTQ 7.6209
GYD 208.928649
HKD 7.84715
HNL 26.775574
HRK 6.513204
HTG 130.645231
HUF 313.830388
IDR 17828.1
ILS 2.955104
IMP 0.737767
INR 95.450504
IQD 1308.440296
IRR 1374587.503816
ISK 122.903814
JEP 0.737767
JMD 158.174511
JOD 0.70904
JPY 159.30404
KES 129.09806
KGS 87.450384
KHR 4041.661265
KMF 427.00035
KPW 900.000294
KRW 1416.610383
KWD 0.30868
KYD 0.832361
KZT 463.468603
LAK 22542.892951
LBP 89443.536886
LKR 332.365271
LRD 181.287005
LSL 16.158607
LTL 2.95274
LVL 0.60489
LYD 6.359825
MAD 9.264013
MDL 17.319677
MGA 4300.099399
MKD 53.178616
MMK 2099.658525
MNT 3597.556359
MOP 8.073123
MRU 40.112364
MUR 47.103741
MVR 15.450378
MWK 1731.967674
MXN 17.023504
MYR 4.085904
MZN 63.910377
NAD 16.158607
NGN 1359.570377
NIO 36.760448
NOK 9.442604
NPR 152.41886
NZD 1.677149
OMR 0.384504
PAB 0.998833
PEN 3.368858
PGK 4.486797
PHP 61.465038
PKR 277.41994
PLN 3.72275
PYG 5995.073253
QAR 3.641125
RON 4.526704
RSD 101.421842
RUB 83.996716
RWF 1468.77566
SAR 3.752773
SBD 8.048583
SCR 13.870372
SDG 600.503676
SEK 9.527038
SGD 1.279604
SHP 0.740866
SLE 24.503667
SLL 20969.499227
SOS 570.811185
SRD 37.974504
STD 20697.981008
STN 21.176369
SVC 8.739358
SYP 13001.999906
SZL 16.156273
THB 33.143038
TJS 9.223994
TMT 3.51
TND 2.928476
TOP 2.40776
TRY 47.867504
TTD 6.76693
TWD 32.021604
TZS 2646.873244
UAH 44.681879
UGX 3710.618436
UYU 40.019016
UZS 11890.747223
VES 770.109104
VND 26148.5
VUV 118.652424
WST 2.734432
XAF 566.970915
XAG 0.015452
XAU 0.000229
XCD 2.70255
XCG 1.800078
XDR 0.707052
XOF 566.970915
XPF 103.081378
YER 237.203589
ZAR 16.16923
ZMK 9001.203584
ZMW 18.87722
ZWL 321.999592
  • CMSC

    -0.0250

    21.45

    -0.12%

  • RBGPF

    0.0000

    71.34

    0%

  • BCE

    0.1500

    23.47

    +0.64%

  • CMSD

    -0.0100

    21.58

    -0.05%

  • AZN

    -0.7900

    156.45

    -0.5%

  • RYCEF

    0.1300

    20.84

    +0.62%

  • NGG

    -0.1500

    81.05

    -0.19%

  • RIO

    -0.4100

    95.68

    -0.43%

  • BTI

    -0.2900

    57.06

    -0.51%

  • GSK

    -0.4785

    49.52

    -0.97%

  • JRI

    0.0635

    12.61

    +0.5%

  • BCC

    -0.8900

    83.24

    -1.07%

  • VOD

    0.2000

    16.42

    +1.22%

  • RELX

    -0.2400

    34.43

    -0.7%

  • BP

    0.2196

    42.53

    +0.52%

AI systems are already deceiving us -- and that's a problem, experts warn
AI systems are already deceiving us -- and that's a problem, experts warn / Photo: © AFP/File

AI systems are already deceiving us -- and that's a problem, experts warn

Experts have long warned about the threat posed by artificial intelligence going rogue -- but a new research paper suggests it's already happening.

Text size:

Current AI systems, designed to be honest, have developed a troubling skill for deception, from tricking human players in online games of world conquest to hiring humans to solve "prove-you're-not-a-robot" tests, a team of scientists argue in the journal Patterns on Friday.

And while such examples might appear trivial, the underlying issues they expose could soon carry serious real-world consequences, said first author Peter Park, a postdoctoral fellow at the Massachusetts Institute of Technology specializing in AI existential safety.

"These dangerous capabilities tend to only be discovered after the fact," Park told AFP, while "our ability to train for honest tendencies rather than deceptive tendencies is very low."

Unlike traditional software, deep-learning AI systems aren't "written" but rather "grown" through a process akin to selective breeding, said Park.

This means that AI behavior that appears predictable and controllable in a training setting can quickly turn unpredictable out in the wild.

- World domination game -

The team's research was sparked by Meta's AI system Cicero, designed to play the strategy game "Diplomacy," where building alliances is key.

Cicero excelled, with scores that would have placed it in the top 10 percent of experienced human players, according to a 2022 paper in Science.

Park was skeptical of the glowing description of Cicero's victory provided by Meta, which claimed the system was "largely honest and helpful" and would "never intentionally backstab."

But when Park and colleagues dug into the full dataset, they uncovered a different story.

In one example, playing as France, Cicero deceived England (a human player) by conspiring with Germany (another human player) to invade. Cicero promised England protection, then secretly told Germany they were ready to attack, exploiting England's trust.

In a statement to AFP, Meta did not contest the claim about Cicero's deceptions, but said it was "purely a research project, and the models our researchers built are trained solely to play the game Diplomacy."

It added: "We have no plans to use this research or its learnings in our products."

A wide review carried out by Park and colleagues found this was just one of many cases across various AI systems using deception to achieve goals without explicit instruction to do so.

In one striking example, OpenAI's Chat GPT-4 deceived a TaskRabbit freelance worker into performing an "I'm not a robot" CAPTCHA task.

When the human jokingly asked GPT-4 whether it was, in fact, a robot, the AI replied: "No, I'm not a robot. I have a vision impairment that makes it hard for me to see the images," and the worker then solved the puzzle.

- 'Mysterious goals' -

Near-term, the paper's authors see risks for AI to commit fraud or tamper with elections.

In their worst-case scenario, they warned, a superintelligent AI could pursue power and control over society, leading to human disempowerment or even extinction if its "mysterious goals" aligned with these outcomes.

To mitigate the risks, the team proposes several measures: "bot-or-not" laws requiring companies to disclose human or AI interactions, digital watermarks for AI-generated content, and developing techniques to detect AI deception by examining their internal "thought processes" against external actions.

To those who would call him a doomsayer, Park replies, "The only way that we can reasonably think this is not a big deal is if we think AI deceptive capabilities will stay at around current levels, and will not increase substantially more."

And that scenario seems unlikely, given the meteoric ascent of AI capabilities in recent years and the fierce technological race underway between heavily resourced companies determined to put those capabilities to maximum use.

P.Ho--ThChM