← Back to all articles

Field note

A bird quiz with working code and incorrect facts

In a trial reported by Stefan Grunert, Qwen3.6-35B-A3B generated a bird quiz from a single request on a MacBook Pro with an M5 Max and 128 GB of memory. The interface and answer evaluation worked, but checking the bird names and descriptions revealed problems. This account concerns one case; runtime measurements and a controlled comparison are not available for this article.

The original request

The following prompt was used for the first version and is reproduced unchanged. It asked for a quiz in a single HTML file that would run in a browser without external libraries, images or internet access.

The quiz was built in the Code area as a new project. There, too, the agent works in a copy: index.html reaches the project folder only after Keep and can then be opened in a browser with a double-click.

Original prompt

Create a single file called index.html: a simple bird quiz that works by opening it in a browser. No external libraries, images or internet access – plain HTML, CSS and JavaScript in one file. Content: 8 common Norwegian garden birds (for example great tit, blue tit, magpie, bullfinch, robin, house sparrow, chaffinch, fieldfare). For each bird, write a short description (2–3 sentences: colours, size, typical behaviour) without mentioning its name. How it works: - Show one description at a time, with 4 answer buttons (the correct bird plus 3 random wrong ones, in random order). - When the player clicks an answer, colour the correct button green and a wrong choice red, and show a short fun fact about the bird. - A "Next" button moves to the next question. - Show progress ("Question 3 of 8") and the current score. - At the end, show the final score with a friendly message and a "Play again" button that reshuffles the questions. Design: calm, nature-inspired colours (greens and earthy tones), large readable text, rounded buttons, centred layout that also works on a phone.

Check function and subject matter separately

Feedback on the first version described a coherent layout and correct highlighting of the answer, including after a wrong selection. It also flagged invented or unsuitable bird names, duplicate answer options and species missing from the requested selection.

These observations concern two different checks. A quiz can process input correctly while its questions or answers are factually wrong. Testing the buttons would not have revealed the content errors. The reported criticisms are not presented here as an independent ornithological assessment.

Revise the content with web search

Stefan subsequently reported that the integrated web search helped correct the errors. Ancilo can search and consult sources in coding sessions as well as chats. A provider must first be configured under System: Wikipedia without an account, or Google through Serper with a personal key.

The Web search switch next to the input decides whether searching is allowed; it applies to chats, tasks and coding alike. When it is off, each query is shown before it is sent and waits for your approval. When it is on, the agent searches on its own. A revision should check not only names but also descriptions, answer assignments and the completeness of the species list against the sources.

Give the correction a defined scope

A further trial can explicitly separate fact checking from interface changes. The following request is a suggested prompt for another attempt, not the wording of the original.

Example request

Check the bird names, descriptions and correct answers in this quiz against traceable web sources. First list questionable statements and their sources. Then correct only the factual content. Preserve the layout and controls. Also check that all requested species are included and each question has exactly one correct answer.

Limits of the account

The case illustrates a possible workflow: try the generated code, check its content separately and use sources to revise questionable statements. It establishes neither a general error rate for the model nor that Norwegian caused the errors. Those conclusions would require comparable tasks, recorded settings and repeated runs.

Web search supplies additional information. Whether a source supports the particular statement, and whether the correction was incorporated accurately, still need checking. Run the quiz again after changing its content to check that the answer text and evaluation logic remain consistent.