Yes, ChatGPT is wrong about your level in the general sense: its estimates are not reliable and tend to be vague, conservative, and poorly calibrated. Your Spanish friend's situation is strong evidence, because a person who scores C1 to C2 in reading should not be labeled B1+ to B2 by a generic chat model. Treat ChatGPT as a conversation tool, not as a certified CEFR test.

ChatGPT has no way to evaluate your skills thoroughly. It cannot observe your listening comprehension, measure your vocabulary recall under pressure, test your pronunciation, or adapt to your best and worst tasks. It also tends to produce middle lane estimates because that is the safest, most general answer for many learners. The CEFR scale describes broad skill sets. A model that guesses from conversation or text will usually avoid claiming a high level, and that can suppress a genuine B2 or C1 learner into a B1-B2 band.

To know your actual level, use multiple sources. First, take a standardized, proctored CEFR test from an established language examination body. These tests are not free, but they are calibrated and accepted by employers, schools, and governments. Do not rely on ChatGPT to generate a mock exam and grade it; the result is not a valid measurement. Second, complete a structured self-assessment against official CEFR can-do statements. For example, can you understand the main ideas of complex texts, write clear detailed texts on a wide range of subjects, and interact with native speakers without obvious effort? If you can do most of these consistently, you are around B2. If you can do them with flexibility, precision, and ease in demanding professional or academic contexts, you are closer to C1.

Third, get feedback from people who are trained to evaluate language. A good teacher or language coach can give you a more accurate profile than ChatGPT. Casual conversation partners can describe whether you seem fluent, but they have no fixed benchmark. Fourth, measure yourself with real tasks. Do a timed writing exercise, listen to a news report or a podcast at native speed, and read a well-known opinion article. Note what caused trouble. This helps you see the gaps in vocabulary, grammar, and comprehension.

Official tests cost money and time, and you may still land between levels because no single test captures everything. Free tools and chat models are convenient but their estimates are not evidence. Conversation practice with friends is valuable but not calibrated. The honest approach is to combine a real test with self-assessment, classroom feedback, and real-world tasks. Doing that will give you a level you can defend, and it will stop you from relying on a model that tells everyone the same safe answer.