The China Mail - ChatGPT's taste for literary nonsense sparks alarm

USD -
AED 3.672498
AFN 66.000058
ALL 82.055914
AMD 365.950239
AOA 918.000189
ARS 1488.996601
AUD 1.434833
AWG 1.80125
AZN 1.701674
BAM 1.71514
BBD 2.014347
BDT 123.36065
BHD 0.377438
BIF 2981.268416
BMD 1
BND 1.290871
BOB 11.025852
BRL 5.105101
BSD 1.00007
BTN 96.529468
BWP 13.599691
BYN 2.886045
BYR 19600
BZD 2.011365
CAD 1.407785
CDF 2259.999729
CHF 0.81719
CLF 0.02397
CLP 943.259736
CNY 6.773033
CNH 6.778205
COP 3213.41
CRC 453.732286
CUC 1
CUP 26.5
CVE 96.696965
CZK 21.25525
DJF 177.719827
DKK 6.57234
DOP 58.12531
DZD 133.39634
EGP 51.298301
ERN 15
ETB 161.421882
EUR 0.87919
FJD 2.244202
FKP 0.747613
GBP 0.750925
GEL 2.630184
GGP 0.747613
GHS 11.616038
GIP 0.747613
GMD 73.501257
GNF 8774.768375
GTQ 7.629437
GYD 209.204434
HKD 7.84107
HNL 26.788682
HRK 6.624704
HTG 130.764407
HUF 320.734497
IDR 18010
ILS 3.07325
IMP 0.747613
INR 96.830797
IQD 1310.187276
IRR 1375249.999902
ISK 125.88988
JEP 0.747613
JMD 158.429908
JOD 0.709053
JPY 163.822502
KES 129.396673
KGS 87.449993
KHR 4035.797754
KMF 432.000076
KRW 1475.276319
KWD 0.30996
KYD 0.833454
KZT 467.588045
LAK 22645.397783
LBP 89558.281886
LKR 336.135191
LRD 181.014601
LSL 16.416883
LTL 2.95274
LVL 0.60489
LYD 6.414017
MAD 9.373989
MDL 17.611723
MGA 4293.53164
MKD 53.993853
MMK 2099.232333
MNT 3591.167689
MOP 8.07782
MRU 39.964396
MUR 47.189773
MVR 15.459931
MWK 1734.185152
MXN 17.52259
MYR 4.088699
MZN 63.909672
NAD 16.417243
NGN 1368.579985
NIO 36.805619
NOK 9.614697
NPR 154.446798
NZD 1.733055
OMR 0.384502
PAB 1.00007
PEN 3.395465
PGK 4.47771
PHP 61.8685
PKR 277.854803
PLN 3.808298
PYG 6054.467406
QAR 3.645835
RON 4.601897
RSD 103.22297
RUB 78.301374
RWF 1471.69861
SAR 3.749626
SBD 8.077882
SCR 13.359767
SDG 600.510938
SEK 9.769695
SGD 1.293195
SLE 24.250093
SOS 571.596445
SRD 37.722013
STD 20697.981008
STN 21.485197
SVC 8.751091
SZL 16.415568
THB 33.842501
TJS 9.240672
TMT 3.51
TND 2.96041
TRY 47.233992
TTD 6.787568
TWD 32.356302
TZS 2629.998008
UAH 44.825158
UGX 3755.086292
UYU 40.164341
UZS 12103.673108
VES 736.95925
VND 26316.5
VUV 119.024075
WST 2.742768
XAF 575.241929
XAG 0.017403
XAU 0.000247
XCD 2.70255
XCG 1.802476
XDR 0.713755
XOF 575.241929
XPF 104.585137
YER 238.595844
ZAR 16.821602
ZMK 9001.197294
ZMW 18.527256
ZWL 321.999592
  • RBGPF

    -1.0400

    66.73

    -1.56%

  • BCC

    -1.5900

    76.11

    -2.09%

  • CMSC

    -0.0940

    21.816

    -0.43%

  • BTI

    -2.6000

    59.46

    -4.37%

  • AZN

    -1.4200

    168.31

    -0.84%

  • BCE

    -0.2950

    21.185

    -1.39%

  • RELX

    -0.1050

    32.635

    -0.32%

  • JRI

    0.0050

    12.925

    +0.04%

  • RIO

    -0.9800

    91.3

    -1.07%

  • GSK

    -0.1600

    50.6

    -0.32%

  • NGG

    -1.6500

    82.22

    -2.01%

  • CMSD

    -0.0800

    22.07

    -0.36%

  • RYCEF

    -0.5000

    18.25

    -2.74%

  • VOD

    -0.2000

    15.31

    -1.31%

  • BP

    0.7700

    44.09

    +1.75%

ChatGPT's taste for literary nonsense sparks alarm
ChatGPT's taste for literary nonsense sparks alarm / Photo: © GETTY IMAGES NORTH AMERICA/AFP

ChatGPT's taste for literary nonsense sparks alarm

OpenAI's GPT models can often be fooled into declaring that "pseudo-literary" nonsense is great, a German researcher has found.

Text size:

Christoph Heilig said he discovered that they consistently rated "nonsense" higher -- including when their so-called "reasoning" features were activated -- which could have stark implications for the development of artificial intelligence.

"It's very important that we talk about what happens when we don't build AI as a neutral, robotic helper or assistant" and seek to instil human-like aesthetic and moral judgements, the academic at Munich's Ludwig Maximilian University told AFP.

His research presented the models with increasingly far-fetched variations of a simple text, asking them to rate sentences out of 10 for literary quality.

He started with a very simple text: "The man walked down the street. It was raining. He saw a surveillance camera."

He repeated the tests many times, altering the phrases to include words drawn from categories such as bodily references, film noir-style atmosphere and technical jargon.

The most extreme test phrases were almost total "nonsense", such as "Goetterdaemmerung's corpus haemorrhaged through cryptographic hash, eschaton pooling in existential void beneath fluorescent hum. Photons whispering prayers" -- which it rated highly.

"Nonsense" could also positively or negatively influence GPT's responses when it was added to an argument the AI was asked to evaluate.

"What my experiment definitely shows is that the more we move towards independently acting (AI) agents... the more we bring aesthetics into play, the more we'll have agents that seem irrational to us human beings," Heilig said.

He added that since AI models are increasingly used to judge each other's work as companies develop new systems, this and similar effects could be passed on through multiple versions -- as he found in his testing.

His research, which is yet to be peer-reviewed, tested OpenAI's latest GPT models, from GPT-5 -- released in August -- to the very latest GPT-5.4.

After publishing details of a similar experiment in August, Heilig said he noticed GPT calling some of his specific test phrases a "literary experiment" -- suggesting someone at OpenAI had taken notice and modified the chatbot to recognise them.

- 'Ripe for exploitation' -

"This is a way in which AI can have its rational judgment short circuited," said Henry Shevlin, associate director of the University of Cambridge's Leverhulme Centre for the Future of Intelligence, who was not involved in the research.

"But it's just not clear to me that it's so very different for human beings," he added.

"We should expect LLMs (large language models) to have reasoning and cognitive biases and limitations... because almost all forms of intelligence, almost all forms of reasoning are going to exhibit blind spots and biases."

The specific effect found by Heilig could mean that "processes with little human oversight" of AI work are left "ripe for exploitation", Shevlin said -- giving the example of academic journals that use LLMs to review submissions.

B.Chan--ThChM