Twitter Proposes; Prolific Disposes - Part I
Using Expected Parrot to test "promising" Ideas with real humans
tl;dr - I had an idea for a <robot> tag that authors could use to identify AI-written passages in blog posts and thought it might be a useful new convention that increased readers’ trust. My simulated agents agreed, but my real human study threw some cold water on it.
The other day, I had an idea for a way that authors could responsibly use AI in their writing. If an author clearly fenced off AI-written text with “robot” tags, perhaps that would solve some of the problems with so-called “AI slop.” Readers would appreciate the candor and increase their trust in the whole thing:
The idea is you’d see passages like this:
<🤖> Sure, I’ll prove the Riemann conjecture for you…</🤖>
I had actually already used this approach in several blog posts. There seems to be a real gap between how much people dislike or claim to dislike “AI slop” and their willingness to read AI-generated text when it’s useful and presented appropriately.
Could the robot tag idea be a paper?!
As the likes and re-tweets rolled in, I started thinking, “Hmm, maybe I could write a paper about this.” Of course, it’s not a “big idea” and the potential upside, research-wise, of introducing a new kind of blog post tag isn’t that high. It’s not the kind of thing I’d sink a ton of time into unless I had some really, really easy way to do the studies…
Let’s use the Expected Parrot Research Agent
But I do have an easy way to do studies! I started up the Expected Parrot Research Agent and described what I wanted to investigate, referencing the original tweet. My full Research Agent transcript is here.
In the transcript, you can see we actually do several different things: explore the idea, play around with framing, develop questions and stimuli, run pilots, revise the study, etc. There’s a part where I tweak the CSS of the “robot” tag in the study to be sent to human participants. One key thing about the Expected Parrot platform is that it’s designed from the ground up for taking a survey given to LLMs and then administering it to real humans.
Collaborator, not a paper factory
If you read the transcript, you’ll see there is a lot of back-and-forth where I propose ideas, critique suggestions, commission pilots, tweak language, etc.:
This is the intended workflow. The Research Agent is not a push button that produces a finished paper. It’s a research collaborator that makes it much faster and cheaper to explore an idea, test assumptions, and improve a study before spending real money on participants.
Promising results but in silico and an underpowered human pilot!
I ran a little pilot with simulated subjects that suggested the idea might be promising. I then ran an underpowered study on Prolific (I’m not made of money). Effects went the “right” way. I could already see the academic paper: maybe not a world-changing piece of scholarship but a nice solid base hit. A contribution squarely in the Information Systems tradition…
Ah crap…
On the basis of my promising pilots, I launched a larger study on Prolific. And when I ran the real study, nothing happened. At all.
Has my robot-tag idea been disproved?
No. I still think it’s a good idea. But maybe not the slam dunk I initially imagined. Maybe I’m now a bit sadder and wiser, but social science studies rarely “prove” anything, one way or another. I could think of ten different reasons why this study doesn’t really speak to the power of the robot tag. But the study is still evidence, and every study provides information that should update our priors on our own imperfect models of the world, even if it doesn’t settle the question. We’re all in this messy business of taking in new data and computing the cross-entropy with respect to the model we already had in our heads. Treating every result as either total confirmation or complete refutation is childish. A more useful question is: How much should this result change what I believe?
That is especially true outside academia, where the goal is often to make a decision rather than publish a paper. Even a well-designed study rarely captures every relevant feature of the real-world problem. But it can still make the decision better.
In this case, the evidence moved me from “This obviously works!” to “This (still) seems plausible but could use more work.” That is definitely not as exciting as proving myself right, but it is probably more useful.
Test your own idea
One reason we built Expected Parrot was to make this kind of small, speculative research project easy to pursue. You can explore an idea, test it with AI, and then run the same study with real people without turning it into a major undertaking.
So, if there is a question you have been curious about, please try it on Expected Parrot and request access to our Research Agent. If you need any help, please get in touch!









