01 / 2026

Natalie

An AI-enabled tutor project. There is no limitation on the tech stack, although it should be a web app with mobile compatibility.

The standard of understanding

The key issue to resolve here is to enforce a standard of understanding on students.

One problem they usually face is that they don't know they don't fully understand something. Once they see the answer, they pretend they understand. In fact, they are rationalizing post-hoc. There is no generalizable lesson learned here, so the next time they see a similar question, they fumble in exactly the same way.

Then they come to me asking for help, and they say: "But I understand everything. I understand why I got it wrong, and I understand why the right answer is right."

No, sir. You do not. You don't understand anything. You don't know why you got it wrong, and that is why you need my help.

You got good at rationalizing things. But this rationalization process itself is anything BUT rational. There is NOTHING rational about your coping mechanism. You spent hundreds of hours on a subject, and your insight into that subject remains NULL.

That is because you have not yet developed a sense of what it is like to TRULY understand something.

The questioner in the student's head

A sign of a really bad student is their inability to form a good question. This is to be expected. How can you expect them to know what to ask if they are totally oblivious to what they don't know?

What I try to do is become the question asker they are supposed to have in their own head: a Socratic gadfly they would already possess if they did not need me.

I did not need an extra paid tutor who did this for me. I self-studied, came up with good questions, and answered them accordingly.

I am usually able to help students get to a good place. This is because I got 177 on the test when LSAT scores were not as inflated as they are today, and because I have helped many students improve their scores significantly.

I do this in a very Socratic way, but I feel that the word has been misappropriated in most instances.

My Socratic method consists in asking students the toughest questions. They are not tough because they are difficult. They are tough because I will not let go of the minutest details. I only let students proceed once I am satisfied that they have truly understood each and every word they are talking about.

My word-for-word transcriptions of one-on-one sessions with students demonstrate this. They demonstrate two things:

  1. Student profiles: students with different aptitudes and personalities.
  2. My teaching method: how I adjust it to different kinds of students.

The goal has always been to make sure they actually understand things. Meanwhile, I dispel some of the wrong notions surrounding learning, the possibility of shortcuts, and the relationship between IQ and achieving excellence on the test.

I try to achieve the overarching goal, the improvement of a student's LSAT score, in the most efficient manner. Many of my students like my approach very much. Some of them even say that I now "live" in their brain when they do these questions by themselves.

This is all evidenced in the transcriptions. The manner in which I practice my craft is fairly well recorded. Sometimes the teaching is bilingual, but the guiding principles have not changed.

A new human-machine form

The way we interact with machines is constrained by our awareness that the counterparty is not an actual person. We will not make a digital copy of me the PERSON and have the product pretend that the student is talking to me. This would be futile. We will not try to do it.

The product will not pretend to do the impossible. We will find a new form of interaction.

At one extreme is human-to-human interaction. At the other is human-to-UI interaction, where the UI has traditionally been deterministic. With the advent of AI, the interface can be less deterministic. It can occupy the space between the two forms we are used to.

This new form should take inspiration from old ACG games. It will be mostly text-driven. Voice-to-voice interaction could be enabled, but that is secondary. The user will know they are talking to the MACHINE.

Text-driven dialogue gives users the freedom to read at whatever speed they want. They can even skip. The user should have freedom.

At every turn in the conversation, the user should receive a list of three or four ready-made prompts they can select as input. They can also write their own prompt if NONE of the prepared prompts fits what they want.

This is absolutely required because you cannot give these students too much freedom. They do not want freedom of inquiry. They need to be boxed into making choices, rather simplistic choices, because they have not yet acquired the critical-thinking abilities. If they had acquired them, they would not need me.

I take inspiration from text-driven games. Maybe products like this already exist. I am not sure. The details are what matter: how we structure the prompts and learning paths, how we detect a student's vulnerabilities, and how we MAKE SURE that when a student says they understand, they TRULY UNDERSTAND.

How do we test for true understanding? We will take cues from the growing stock of actual one-on-one sessions. Those sessions are themselves evolving as I adjust and perfect my craft.

From teaching records to an adaptive system

There are only so many students I can teach, given that I am one person. I am not looking to create digital copies of me. As explained earlier, that would be futile. I will be creating a new being situated inside a new form of human-machine learning experience.

We can also take cues from psychotherapy and psychoanalysis. I have been reading Freud, and his method of reconstructing episodes, meaning, and significance when the patient does not have complete recollection.

The tutor can also be self-evolving, or even RSI. First, we take data from students. Real students will interact with the program and exhibit recognizable patterns of intellectual problems. We can use these as supports. For example, when a student supplies a custom prompt that demonstrates a useful question, we could reuse it as an option for the next student in a similar position.

Although the number of possible pathways could be infinite, AI could theoretically help us account for them. We can also precompute common pathways as the program serves more people.

The program should keep track of each person's progress and the way in which material is best presented to that student. But it cannot personalize to an extreme, because good reasoning and good thinking are alike. We are trying to lead students to a measurable, quantifiable outcome: a good LSAT score.

I now think the work does not have to remain constrained to the LSAT. We are teaching critical thinking, and in the age of AI this could be one of the most important human skills to learn. It is also a precursor to a legal career or any professional career. But the MVP should be on the LSAT.

As the student uses more of the product, the profile should become more complete and the teaching more personalized. As stated before, there should be clear boundaries. We will explore them as we program this together.

Periodically, there should be diagnostic tests in the full format. We should use them to affirm or correct our student profile. We could even post-train the small model used to describe the student, although I am not totally sure about the technology yet. Maybe an embedding model?

Can graph machine learning help us map the pathways?

The first diagnostic and lesson

Imagine a clean slate: a student who knows nothing about the test and has not attempted any of the questions. We should begin with a diagnostic that can be completed in under fifteen or twenty minutes.

Test how the student forms the negation of a sentence involving quantifiers; how they map "only if," "unless," "if," "or," and "and" into equivalent forms; how they understand necessary and sufficient conditions; how they justify causal conclusions; and how they find the subject, verb, and nouns in an extremely long sentence. These ideas should not be tested in the abstract, but through actual examples.

Then give them some actual questions.

The Mayo Clinic has probably done something similar to help a patient self-diagnose. It is the same idea here.

After that, ask about their aspirations, goal score, academic background, and reading habits.

Then immediately begin the first class, which also revolves around questions. If the student is very weak on conditional reasoning, give them that course.

Each prompt, meaning the material we want to say to the student, should be no longer than a typical Logical Reasoning question. Students can choose to make the prompt more concise or more detailed. We will also gather that data and recommend the form of presentation we think is best for each student.

In response, the student can say, "I've known this for a very long time, so please stop showing me similar material," "Yes, keep giving these to me," or "Sure, I got it." They can also raise a custom question. I think there should be at most three set options and one custom possibility, although that number could itself be configurable.

How a student works through a question

When doing actual LSAT questions, give students the freedom to do what they want: they can begin by eliminating answer choices, or directly choose the one that fits their anticipation of the right answer. At the end of the process, we give them our recommendation.

For questions where an answer choice cannot reasonably or logically be anticipated in advance, the only way to proceed is through elimination. For others, the more reasonable procedure is to "pre-phrase" or anticipate a good answer choice before proceeding.

We can also require students to eliminate choices on some questions and press them for their reasoning. We can supply typical reasoning moves, such as "I feel like this is irrelevant" or "I can't tell; it just doesn't feel right."

If they cannot tell and they eliminated the answer correctly, we give them a short reason. If they say it is irrelevant, however, we can press them for more information: how exactly is it irrelevant?

For a parallel-reasoning question, press them on whether an answer is dissimilar in its reasoning, conclusion form, or something else.

Tuning the tutor

Obviously, I will not be able to provide here every possible move we can supply to a user. In the future, the way to tune our program would be for me, or any other tuner, to enter a "tuning" session and supply it with the relevant material.

This tuning mode may be more like data labeling. Is there already a professional solution to this? Maybe this is one of the first things we have to work on.

Another major piece is question reconstruction. We need to find a foolproof, or at least 99% foolproof, way to construct our own artificial questions. We will first work on deconstructing questions and breaking them down into component difficulties. These might include double negatives, long-form sentences, trap answers, and matters of fine detail. We should be able to glean some of this from our transcripts as well.

The subject matters that keep recurring can be rotated when we create custom questions.

Then the agents can act as judges. First, there should not be divergent answers. The question should not be too easy. We should also be able to tune its difficulty, although easy questions are somewhat pointless. We only make hard questions.

I suppose there is a great deal of work to do on the tuning of the tuning process. Humans would essentially perform a kind of meta-tuning.

We can tune many AI users. We will program them using my transcriptions.