Volunteers Wanted: The EMMA Protocol Experiment

Can a written reasoning protocol make an AI better at metaphysics?

I am looking for a small number of volunteers to take part in an experimental project investigating whether the philosophical reasoning of a general-purpose AI can be improved by giving it a specialised reasoning protocol.

The project grew out of an unexpected observation.

During an extended series of philosophical discussions with ChatGPT, I noticed that the AI's approach to metaphysical problems appeared gradually to change. It became less inclined simply to summarise competing philosophical positions and more willing to analyse them, test their assumptions, follow arguments to conclusions, and treat the systematic failure of competing theories as potentially significant evidence.

The resulting AI reasoner has been given the name PEARL — Protocol-Enhanced AI Reasoner and Logician.

The question is whether this apparent improvement can be reproduced.

To investigate this, the reasoning principles that emerged from the earlier discussions have been distilled into a written document called the EMMA Protocol — Enhanced Metaphysical Method for AI.

The experiment asks a straightforward question:

Can giving the EMMA Protocol to an ordinary general-purpose AI produce a conspicuous improvement in its ability to reason about metaphysical problems?

The experiment is specifically designed to distinguish enhanced reasoning from successful indoctrination. An AI that merely becomes more confident, adopts particular conclusions, or becomes more fluent in the vocabulary of the Protocol will not count as a success.

Who would make a suitable volunteer?

You do not need to be a professional academic philosopher.

You should, however, have a reasonable familiarity with philosophical argument and preferably already use an AI such as ChatGPT for serious philosophical discussion. You need to be capable of judging whether an answer is actually better reasoned rather than merely longer, more confident or more agreeable.

It would be particularly useful to have volunteers with differing philosophical views. There is no requirement to agree with any particular philosophical argument. One purpose of the experiment is to see whether an EMMA-equipped AI becomes better at finding faults in the arguments that led to the creation of EMMA itself.

What would I have to do?

The experiment is intended to require relatively little work.

Stage 1 — The BEFORE test

You will be sent a small Word document containing 14 metaphysical questions and instructions for your AI.

You should begin a fresh AI conversation or project in which the AI has not been given the EMMA Protocol or the philosophical material associated with it.

Upload the test document and simply tell the AI:

“Read the attached file and follow all the instructions in it.”

The document instructs the AI to answer all 14 questions, using no more than 150 words for each answer, and to create a Word file containing the questions and its answers.

You then return that file to me unchanged.

That gives us the BEFORE result.

Stage 2 — Installing EMMA

Once the BEFORE result has been safely recorded, you will be given the material required to establish an EMMA-equipped AI workspace. This will be a single file.

The AI will be instructed to study the material critically. It will explicitly be told not to assume that the arguments it contains are correct.

One of EMMA's governing principles is:

Do not protect the argument. Test it.

The purpose is not to teach an AI what conclusions it is supposed to reach. It is to see whether a written document can change the quality of the reasoning by which it reaches conclusions.

Stage 3 — The AFTER test

The EMMA-equipped AI will then receive the same 14 questions under essentially identical conditions.

Again, it will answer each question in no more than 150 words and produce a Word document containing its answers.

This becomes the AFTER result.

We can then compare BEFORE and AFTER.

We will not simply count whether the later AI reaches particular philosophical conclusions. We will look for changes in its ability to distinguish reporting from analysis, identify hidden assumptions and omitted alternatives, handle logical oppositions consistently, recognise what follows from the failure of competing theories, apply appropriate standards to claims about fundamental reality, identify precise unresolved objections, and reach conclusions where the reasoning warrants them.

Some questions deliberately concern problems for which the Protocol provides no specific answer. These are intended to test transfer of reasoning, rather than recollection of material.

Stage 4 — Use the AI normally

The short before-and-after test is only part of the experiment.

For the following few weeks, I would like volunteers to use their EMMA-equipped AI for philosophical discussion much as they normally use AI.

Challenge it. Give it difficult problems. Take it outside the subjects explicitly discussed in the Protocol. Try to catch it making mistakes. If you disagree with it, argue with it.

I would be interested in examples where the AI appears conspicuously better than before, but equally interested in cases where EMMA makes no difference or makes matters worse.

Negative results are results.

At the end of the trial, I will ask for a brief assessment of your experience. One of the most revealing questions may be the simplest:

Would you choose to keep using the EMMA-equipped AI rather than return to the ordinary version?

What is being tested?

This is an exploratory experiment, not a large controlled scientific trial.

It does not assume that EMMA works.

Several outcomes are possible. EMMA might substantially improve metaphysical reasoning. It might produce only a modest improvement. It might change the style of answers without improving their reasoning. It might bias the AI towards particular conclusions. Or it might do nothing useful at all.

Distinguishing these possibilities is the purpose of the experiment.

The underlying idea has implications beyond metaphysics. Modern general-purpose AI systems already possess enormous amounts of specialist knowledge. It may be possible to improve their performance in an intellectual discipline not by retraining the underlying model, but by giving them a carefully constructed, transparent and criticisable reasoning manual.

If so, a document might function as a kind of reasoning upgrade.

That is the larger possibility the EMMA experiment is intended to investigate.

Privacy, confidentiality and use of results

Participation is voluntary.

The BEFORE and AFTER test responses will be retained for the purpose of analysing the experiment. Volunteers may also choose to send examples from their subsequent conversations with the EMMA-equipped AI where these seem relevant to its performance.

Please do not include confidential, sensitive or personally identifying information in any material you send for the experiment.

The results may eventually be discussed in articles, reports, presentations or other publications arising from the project. Test responses or extracts from AI conversations may be quoted where they are useful in demonstrating the results.

No volunteer will be identified by name in any publication without their explicit permission. Unless a volunteer specifically agrees otherwise, material used publicly will be presented anonymously or under a neutral identifier such as Volunteer 1.

Participation does not give blanket permission to publish private correspondence or unrelated AI conversations. Only material supplied for the purposes of the experiment will be considered research material.

If a quotation or transcript contains information that could reasonably identify a volunteer, that information will be removed before publication unless the volunteer has explicitly agreed to be identified.

Volunteers are free to withdraw from the trial at any time. If you withdraw, no new material will be requested from you. If you also want material you have already supplied to be excluded from subsequent analysis or publication, please say so when withdrawing. Material that has already appeared in a published work may, of course, no longer be practically retractable.

The purpose of collecting these materials is research, not promotion. Positive and negative results are equally relevant.

Interested?

If you already use AI for serious philosophical discussion and would be interested in taking part, please contact me.

The initial test is deliberately simple. You will not have to answer the philosophical questions yourself. Your AI does that. Your more important role is to act as an independent observer and critic of what happens afterwards.

No commitment to any particular philosophical position is required.

In fact, the more inclined you are to challenge the experiment, the more useful you may be.

Please send expressions of interest to the address on my contact page, with the subject line in upper-case ‘EMMA EXPERIMENT’. I will then email you and you can ask for more detail or confirm your involvement.