The China Mail - Inbred, gibberish or just MAD? Warnings rise about AI models

USD -
AED 3.672504
AFN 65.503991
ALL 80.193613
AMD 365.443623
ANG 1.789783
AOA 918.000367
ARS 1475.150612
AUD 1.415829
AWG 1.80125
AZN 1.70397
BAM 1.690479
BBD 2.011669
BDT 122.606854
BGN 1.696366
BHD 0.37668
BIF 2985.954449
BMD 1
BND 1.277583
BOB 11.640953
BRL 5.222404
BSD 0.998833
BTN 95.261679
BWP 13.45607
BYN 3.037988
BYR 19600
BZD 2.008816
CAD 1.38765
CDF 2273.000362
CHF 0.813512
CLF 0.023251
CLP 913.600415
CNY 6.743204
CNH 6.74452
COP 3115.519253
CRC 449.371191
CUC 1
CUP 26.5
CVE 95.306625
CZK 20.930304
DJF 177.864212
DKK 6.460604
DOP 58.464065
DZD 131.663425
EGP 49.854358
ERN 15
ETB 161.573793
EUR 0.864504
FJD 2.234204
FKP 0.737767
GBP 0.739153
GEL 2.610391
GGP 0.737767
GHS 10.937378
GIP 0.737767
GMD 73.503851
GNF 8773.931458
GTQ 7.6209
GYD 208.928649
HKD 7.84715
HNL 26.775574
HRK 6.513204
HTG 130.645231
HUF 313.830388
IDR 17828.1
ILS 2.955104
IMP 0.737767
INR 95.450504
IQD 1308.440296
IRR 1374587.503816
ISK 122.903814
JEP 0.737767
JMD 158.174511
JOD 0.70904
JPY 159.30404
KES 129.09806
KGS 87.450384
KHR 4041.661265
KMF 427.00035
KPW 900.000294
KRW 1416.610383
KWD 0.30868
KYD 0.832361
KZT 463.468603
LAK 22542.892951
LBP 89443.536886
LKR 332.365271
LRD 181.287005
LSL 16.158607
LTL 2.95274
LVL 0.60489
LYD 6.359825
MAD 9.264013
MDL 17.319677
MGA 4300.099399
MKD 53.182938
MMK 2099.658525
MNT 3597.556359
MOP 8.073123
MRU 40.112364
MUR 47.103741
MVR 15.450378
MWK 1731.967674
MXN 17.023504
MYR 4.085904
MZN 63.910377
NAD 16.158607
NGN 1359.570377
NIO 36.760448
NOK 9.442604
NPR 152.41886
NZD 1.697217
OMR 0.384504
PAB 0.998833
PEN 3.368858
PGK 4.486797
PHP 61.465038
PKR 277.41994
PLN 3.72275
PYG 5995.073253
QAR 3.641125
RON 4.526704
RSD 101.421842
RUB 84.182981
RWF 1468.77566
SAR 3.752773
SBD 8.048583
SCR 13.755996
SDG 600.503676
SEK 9.527038
SGD 1.279604
SHP 0.740866
SLE 24.503667
SLL 20969.499227
SOS 570.811185
SRD 37.974504
STD 20697.981008
STN 21.176369
SVC 8.739358
SYP 13001.999906
SZL 16.156273
THB 33.143038
TJS 9.223994
TMT 3.51
TND 2.928476
TOP 2.40776
TRY 47.867504
TTD 6.76693
TWD 32.021604
TZS 2646.873244
UAH 44.681879
UGX 3710.618436
UYU 40.019016
UZS 11890.747223
VES 770.109104
VND 26148.5
VUV 118.652424
WST 2.734432
XAF 566.970915
XAG 0.015452
XAU 0.000229
XCD 2.70255
XCG 1.800078
XDR 0.707052
XOF 566.970915
XPF 103.081378
YER 237.203589
ZAR 16.16923
ZMK 9001.203584
ZMW 18.87722
ZWL 321.999592
  • CMSC

    -0.0250

    21.45

    -0.12%

  • CMSD

    -0.0100

    21.58

    -0.05%

  • NGG

    -0.1500

    81.05

    -0.19%

  • RELX

    -0.2400

    34.43

    -0.7%

  • BCC

    -0.8900

    83.24

    -1.07%

  • RBGPF

    0.0000

    71.34

    0%

  • RIO

    -0.4100

    95.68

    -0.43%

  • GSK

    -0.4785

    49.52

    -0.97%

  • BCE

    0.1500

    23.47

    +0.64%

  • AZN

    -0.7900

    156.45

    -0.5%

  • BTI

    -0.2900

    57.06

    -0.51%

  • JRI

    0.0635

    12.61

    +0.5%

  • VOD

    0.2000

    16.42

    +1.22%

  • RYCEF

    0.1300

    20.84

    +0.62%

  • BP

    0.2196

    42.53

    +0.52%

Inbred, gibberish or just MAD? Warnings rise about AI models
Inbred, gibberish or just MAD? Warnings rise about AI models / Photo: © AFP/File

Inbred, gibberish or just MAD? Warnings rise about AI models

When academic Jathan Sadowski reached for an analogy last year to describe how AI programs decay, he landed on the term "Habsburg AI".

Text size:

The Habsburgs were one of Europe's most powerful royal houses, but entire sections of their family line collapsed after centuries of inbreeding.

Recent studies have shown how AI programs underpinning products like ChatGPT go through a similar collapse when they are repeatedly fed their own data.

"I think the term Habsburg AI has aged very well," Sadowski told AFP, saying his coinage had "only become more relevant for how we think about AI systems".

The ultimate concern is that AI-generated content could take over the web, which could in turn render chatbots and image generators useless and throw a trillion-dollar industry into a tailspin.

But other experts argue that the problem is overstated, or can be fixed.

And many companies are enthusiastic about using what they call synthetic data to train AI programs. This artificially generated data is used to augment or replace real-world data. It is cheaper than human-created content but more predictable.

"The open question for researchers and companies building AI systems is: how much synthetic data is too much," said Sadowski, lecturer in emerging technologies at Australia's Monash University.

- 'Mad cow disease' -

Training AI programs, known in the industry as large language models (LLMs), involves scraping vast quantities of text or images from the internet.

This information is broken into trillions of tiny machine-readable chunks, known as tokens.

When asked a question, a program like ChatGPT selects and assembles tokens in a way that its training data tells it is the most likely sequence to fit with the query.

But even the best AI tools generate falsehoods and nonsense, and critics have long expressed concern about what would happen if a model was fed on its own outputs.

In late July, a paper in the journal Nature titled "AI models collapse when trained on recursively generated data" proved a lightning rod for discussion.

The authors described how models quickly discarded rarer elements in their original dataset and, as Nature reported, outputs degenerated into "gibberish".

A week later, researchers from Rice and Stanford universities published a paper titled "Self-consuming generative models go MAD" that reached a similar conclusion.

They tested image-generating AI programs and showed that outputs become more generic and strafed with undesirable elements as they added AI-generated data to the underlying model.

They labelled model collapse "Model Autophagy Disorder" (MAD) and compared it to mad cow disease, a fatal illness caused by feeding the remnants of dead cows to other cows.

- 'Doomsday scenario' -

These researchers worry that AI-generated text, images and video are clearing the web of usable human-made data.

"One doomsday scenario is that if left uncontrolled for many generations, MAD could poison the data quality and diversity of the entire internet," one of the Rice University authors, Richard Baraniuk, said in a statement.

However, industry figures are unfazed.

Anthropic and Hugging Face, two leaders in the field who pride themselves on taking an ethical approach to the technology, both told AFP they used AI-generated data to fine-tune or filter their datasets.

Anton Lozhkov, machine learning engineer at Hugging Face, said the Nature paper gave an interesting theoretical perspective but its disaster scenario was not realistic.

"Training on multiple rounds of synthetic data is simply not done in reality," he said.

However, he said researchers were just as frustrated as everyone else with the state of the internet.

"A large part of the internet is trash," he said, adding that Hugging Face already made huge efforts to clean data -- sometimes jettisoning as much as 90 percent.

He hoped that web users would help clear up the internet by simply not engaging with generated content.

"I strongly believe that humans will see the effects and catch generated data way before models will," he said.

C.Mak--ThChM