After an observed teaching practice lesson, there is a moment that most teacher trainers will recognise. The trainee has taught. We have taken pages of notes. There will be feedback, discussion, reflection and action points. And somewhere in that process comes the judgement: to standard. Or, sometimes, not to standard. At some centres I have worked with, those are effectively the only two judgements trainees hear after individual teaching practice lessons. At others, tutors make finer distinctions: not to standard, to standard – weak, to standard, to standard – strong, perhaps even above standard. I have been wondering recently which approach actually serves trainees better. I do not think there is an obvious answer.

The attraction of keeping it simple

There are good reasons for using only to standard and not to standard. CELTA is continuously assessed. Individual teaching practice lessons are not mini examinations producing a series of grades that can simply be averaged at the end. Cambridge states that teaching practice assessment is based on the candidate’s overall performance across the course, and the eventual grades are Pass A, Pass B, Pass or Fail. So perhaps attaching increasingly fine judgements to individual lessons risks creating a false sense of mathematical precision. A trainee who receives above standard in TP3 might understandably begin translating that into, “I’m heading for a Pass A.” But that conclusion may not be justified. Development on an initial teacher-training course is rarely linear. Someone can perform very strongly in one lesson and struggle badly in the next. Different lesson types expose different strengths and weaknesses. Planning, language awareness, classroom management and responsiveness to learners do not necessarily develop at the same rate.

There is also a motivational argument for keeping the categories broad. Imagine a trainee receiving to standard – weak after working extremely hard on a lesson. We may intend the distinction to communicate, “You have met the required standard, but there are areas that need attention.” What they may hear instead is, “You nearly failed.” The label can become much louder than the feedback around it. Research on feedback gives us some reason to take this seriously. Kluger and DeNisi’s well-known meta-analysis found that feedback does not automatically improve performance. In fact, more than a third of the feedback interventions they examined reduced it. One explanation they proposed was that feedback becomes less effective when it moves attention away from the task and towards the self. That feels highly relevant to teacher training. If a trainee leaves feedback thinking mainly, “I’m weak”, rather than, “I need to reduce my instructions, monitor more actively and check whether learners have understood the target language”, we may have given them more information but less useful feedback. There is something appealing, then, about saying simply: You are currently meeting the standard. Now let’s talk about your teaching.

But what gets lost?

The problem is that to standard covers an enormous amount of territory. Consider two trainees. One has just produced a lesson that met the criteria, but only just. The aims were broadly achieved, yet instructions caused confusion, monitoring was inconsistent and the final practice stage was rushed. Another trainee also receives to standard. Their lesson was confident, responsive and well staged. There were some issues, but they are already demonstrating several behaviours we would normally associate with considerably stronger performance. Technically, both judgements may be accurate. Educationally, though, are we giving both trainees enough information? This is where the argument for more differentiated judgements becomes interesting.

Butler and Winne describe feedback as central to self-regulated learning. Learners need to compare where they currently are with where they are trying to go. That comparison helps them decide what to do next. For a trainee teacher, knowing “I meet the standard” gives one piece of information. Knowing “I meet it, but I am currently close to its lower boundary” gives another. The second may generate a very different response. A trainee repeatedly told that their lessons are to standard may reasonably assume that things are progressing comfortably. Detailed written and oral feedback might tell another story, but labels are powerful. People use them as shortcuts. If tutors are privately thinking, “Yes, this is to standard, but it is becoming borderline and we need to see stronger progress,” while the trainee hears only, “To standard,” there is a possible information gap. The trainee cannot respond to information they have not been given.

And what about strong trainees?

The same issue exists at the other end. Some centres are understandably cautious about saying above standard. There is a fear that a trainee may become complacent: “I’m already above standard, so I’m obviously doing fine.” There may also be concern about creating expectations around final grades which tutors cannot responsibly guarantee halfway through a course. Both concerns make sense. Yet withholding information has consequences too. If someone is doing exceptionally well, should they know that?

Ryan and Deci’s work on self-determination theory suggests that feedback which supports a person’s sense of competence can strengthen motivation, particularly when it is informational rather than controlling. Perhaps telling somebody that their teaching is currently very strong does not inevitably make them complacent. It may do the opposite. A trainee might hear, “This was above the expected standard at this stage. Now the challenge is to sustain that quality while becoming more flexible and responsive.” That is not the same as saying, “Congratulations. Pass A secured.” The distinction seems important.

Maybe the question is not how many labels we use

This is where I keep going backwards and forwards. A five-point system appears to provide more precision: not to standard, to standard – weak, to standard, to standard – strong, above standard. But precision can be deceptive. How reliably can different tutors distinguish to standard from to standard – strong? Would another tutor make exactly the same judgement? At what point does a useful developmental distinction start looking like a grading scale that was never designed to be one? A two-point system avoids some of those problems, but simplicity can also conceal meaningful differences. To standard may be technically correct while still leaving the trainee unsure whether they are thriving, progressing adequately or only just keeping their head above water.

And this brings me back to something Shute argues in her review of formative feedback: useful feedback needs to be specific enough to help the learner change subsequent performance. Perhaps the most interesting question, then, is not simply whether we should have two categories or five. It is what the trainee knows after receiving our judgement that they did not know before. Do they know how securely they are meeting the criteria? Do they understand whether their development is moving in the right direction? Do they know which aspects of their teaching are becoming strengths, which weaknesses are becoming serious, and what they should try in the next lesson? And perhaps most importantly, does our feedback make them want to act on that information?

Carless and Boud use the term feedback literacy to describe learners’ capacity to understand feedback, make judgements about their own work and take action as a result. That idea seems particularly appropriate for teacher education. Ultimately, we do not want trainees to become dependent on tutors announcing whether their latest lesson was weak, strong or somewhere in between. We want them gradually to become better at making those judgements themselves. But while they are developing that ability, how much should we tell them?

I am still not convinced that either approach has a decisive advantage. Too many categories can turn developmental feedback into a hunt for grades. Too few can deprive trainees of information they may genuinely need. Perhaps a trainee who hears to standard – weak becomes discouraged. Perhaps another finally understands the urgency of making a change. Perhaps above standard makes one trainee complacent. Perhaps it gives another trainee the confidence to experiment, take risks and develop further. That is why I find the question more interesting than the answer.

When we assess teaching practice, how much should trainees know about exactly where within the standard we think they are?

I would be particularly interested to hear how other trainers and centres approach this.

References

Black, P. & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7–74.

Butler, D. L. & Winne, P. H. (1995). Feedback and self-regulated learning: A theoretical synthesis. Review of Educational Research, 65(3), 245–281.

Carless, D. & Boud, D. (2018). The development of student feedback literacy: enabling uptake of feedback. Assessment & Evaluation in Higher Education, 43(8), 1315–1325.

Kluger, A. N. & DeNisi, A. (1996). The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory. Psychological Bulletin, 119(2), 254–284.

Ryan, R. M. & Deci, E. L. (2000). Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. American Psychologist, 55(1), 68–78.

Shute, V. J. (2008). Focus on formative feedback. Review of Educational Research, 78(1), 153–189.

Posted in

Leave a comment