This article is part of our comprehensive guide to hypermobility and Ehlers-Danlos syndrome.
The Beighton score is nine points built out of five movements, and somebody can run it on you in roughly the time it takes to find a chair [1][2]. It was designed to count hypermobile joints across large groups of people, and for that particular job it’s a genuinely sensible bit of kit, as it needs almost no equipment, it’s quick, and two different examiners will usually land on the same number [3][4]. What it was never built to do though, is decide whether anybody takes your symptoms seriously, and that, unfortunately, is a fair description of what it does now [5].
So, if you’ve arrived here with a number already sitting in your head, either a 2 that got you shown the door or a 9 that got you a shrug and a leaflet, the honest answer runs much the same in both directions, as that number describes how far five movements went, on one day, in one room, with one person’s hands on you [3][6]. It doesn’t grade severity, and it knows nothing at all about your pain, your subluxations, your gut, your fatigue, or the fact that you could fold flat at eleven and can’t get past your shins now. All of those things matter a great deal to whether you have a hypermobility disorder, and not one of them is in the nine points [7][6].
None of which makes the score rubbish, by the way. It’s the one measure of joint laxity that nearly everybody, everywhere, more or less agrees on, and a shared language is worth a lot in a field this messy [5][7], so the trouble only really starts the moment a shared shorthand gets treated as a verdict.
This article covers:
ToggleWhat the Nine Points Actually Are
Five movements, scored on both sides where you’ve got two of something, one point each, and nine at the top [1][2].
– The little fingers: with the hand relaxed and somebody else doing the bending, the fifth finger goes back past 90 degrees. One point per hand, so two available [8][1].
– The thumbs: with the wrist bent forward, the thumb can be drawn down far enough to touch the forearm. One point per thumb, so another two [8][9].
– The elbows: the elbow travels back past straight by more than 10 degrees. One point per elbow [1].
– The knees: the same idea at the other end, where the knee goes back past straight by more than 10 degrees. One point per knee [1].
– The trunk: standing with your knees straight, you bend forward and put both palms flat on the floor. One point, and that’s the ninth [2][1].
That’s the whole of the famous test, and it’s worth seeing it written out plainly, because a nine point score sounds like a thorough piece of measurement right up until you count the joints involved and notice how few of them there actually are. Four of the nine points are in your hands, four more are elbows and knees, and the last one is a hamstring test as much as it’s anything else.
Two words in there are doing a great deal more work than they look like they’re doing. The first is passive, as the finger and the thumb are supposed to be moved by the examiner rather than by you, which sounds like a technicality and really isn’t, given that whether the test gets done passively or actively changes what comes out of it. That, along with whether the elbows and knees are measured with a goniometer or just eyeballed, and what starting position everybody is working from, sits in a small pile of procedural questions that have never properly been settled [3][8]. So the number on your letter is a fair bit softer than a number on a letter usually feels.
The second word is today, as the score measures what your joints did during that appointment, rather than what they did when you were nine, or what they’ll do after a winter of flares and guarding, or what they do on the days you can barely get off the sofa.
There’s also nothing in between. Each movement is a yes or a no, so a knee that goes back a little past straight scores exactly the same as one that looks like it’s been put on backwards, and an elbow one degree short of the threshold scores nothing at all. There are no gears on it, it’s nine switches, and a switch can’t tell you how much.
Can you score yourself. Sort of, and it’s worth knowing where the rough edges are. Two of the five movements are meant to be done to you rather than by you, so a fifth finger you’ve bent back with your other hand is not quite the same measurement as one bent back by somebody else, and thumb to forearm has exactly the same issue. The elbows and knees are genuinely hard to judge on yourself without a second pair of eyes, as ten degrees past straight is a lot less obvious than it sounds and a photograph from the side is more honest than a mirror (which is why half the photos in every hypermobility group are somebody’s elbow, taken side on, with a caption asking whether that counts). The trunk one is the only one you can really do alone, and it leans on your hamstrings as much as anything else.
For what it’s worth, self reported versions of the score do hold up reasonably well in the right setting, including formats built around line drawings of each movement, so scoring yourself isn’t a waste of time [7]. Treat the result as a rough description to bring to an appointment rather than as the number, and don’t get attached to a single point either way, as a single point is precisely the sort of thing two examiners disagree about.
Now, the longer list is what the score doesn’t do! It isn’t a symptom scale, it doesn’t grade how bad anything is, and it carries no information at all about pain, fatigue, how often things come out of joint, how well you manage day to day, or anything going on in your gut, your skin or your blood pressure [3][7][6]. It also doesn’t diagnose hEDS or hypermobility spectrum disorder by itself, and it was never designed to [6].
That last part gets lost constantly, and you can see why, as a number feels like an answer. Nine points, a threshold, in or out, done. Except the thing being counted is joint range, and the thing you actually want to know about is a disorder, and those are two different objects with an awful lot of complicated territory sitting in between them.
Reliable and Valid Are Not the Same Thing
Two words get used here as though they mean the same thing, and they really don’t.
Reliability means that if I test you and then somebody else tests you, we get the same answer. Validity means the answer we agree on is genuinely measuring the thing we claimed to be measuring. The Beighton score does well on the first one and considerably less well on the second, and the space between those two words is where nearly all of the argument lives.
The reliability really is good, particularly where a structured method is used, and it holds up both between different examiners and for the same examiner testing twice [3][4]. That isn’t a small thing either, as a lot of clinical measures fall apart the moment a second pair of hands has a go, and this one mostly doesn’t.
But reliability is a statement about the ruler, not about what’s being measured with it. Two people can agree beautifully on a number that isn’t capturing what it says on the tin, and that’s roughly the position here: the consistency is real, and it doesn’t rescue the validity problem sitting underneath it [3][5].
The validity problem is simply that the score samples a handful of joints, leans heavily towards the upper limb, and can miss hypermobility at the shoulders, the hips, the ankles and other places that matter clinically, which means it doesn’t cleanly capture whole body hypermobility even though whole body hypermobility is exactly what it gets used to establish [5][10][11].
So why is it everywhere. Partly history and partly practicality, and the practicality is not to be sniffed at. It’s standardised, it’s quick, it needs no kit, and because everybody uses it, a score of 6 recorded in one country means something recognisable to somebody reading about it in another. That shared vocabulary is genuinely why it remains the dominant common language in hypermobility work, across research and guidance both [5][7][12].
A shared language is useful, it’s just not the same thing as a correct one, and the field has been quietly living with that gap for decades.
Why It Ended Up Carrying This Much Weight
None of this was anybody’s plan, by the way, which does change how you read the thing.
The score started life as a way of counting joint laxity across a whole population, and the early work was exactly that, measuring articular mobility across large groups of people to find out how common hypermobility actually was [13]. That’s a perfectly respectable use for a fast, crude, reproducible measure. When you’re describing a population, you don’t really need to be right about any individual, you need to be consistently wrong in the same direction, and the Beighton score is superb at that.
Then it got adopted, as it’s practical, it’s standardised, it needs no equipment, and it became the shared vocabulary across research and guidance both, which is a genuine achievement and the reason it’s still the dominant measure [5][7][12]. Once everybody was using the same nine points, the nine points became the definition of the thing rather than one estimate of it, and the estimate quietly got promoted.
And from there it slid into being a gate, which is the step that did the damage. A measure built for counting populations ended up deciding individual access to assessment, and the reviews now say plainly that this is the wrong job for it: not the principal tool for distinguishing localised from generalised hypermobility, and not something to use alone to rule generalised hypermobility out [5][12].
So when a nine point score decides whether you get a referral, the problem isn’t somebody misusing a diagnostic instrument, it’s that there was never a diagnostic instrument in the first place, just a population counting tool that got handed a job nobody designed it for.

Where the Score Actually Sits in a Diagnosis
Before the cutoffs make any sense, the words need sorting out, as four or five different things get squashed into the phrase “being hypermobile” and the score only speaks to one of them.
Joint hypermobility on its own is a physical trait. Some people’s joints move further than most people’s, that’s all it means, and on its own it’s neither a diagnosis nor a problem. Plenty of people have it, know it, and are perfectly well. The reviews are careful to keep that separate from the symptomatic conditions, which need symptoms, features across more than one body system, other explanations ruled out, and somebody exercising clinical judgement before anybody gets anywhere near a label [14][15][16].
The framework brought in during 2017 was an attempt to put that distinction on a proper footing. It pulled apart asymptomatic joint hypermobility, the symptomatic spectrum disorders, and the syndromic conditions, and it retired the older labels, so benign joint hypermobility syndrome and joint hypermobility syndrome gave way to the hypermobility spectrum disorders and a tighter definition of hEDS [17][14]. The word benign was doing nobody any favours at all, so that one was well overdue.
It also broke hypermobility itself into patterns rather than a single yes or no, and these four are worth knowing by name, as which one you’ve got changes what your score will look like:
– Generalised hypermobility: the whole body pattern, more than one region involved, and the thing the Beighton score is nominally trying to detect [17].
– Localised hypermobility: one joint, or one region of the body, moving further than it should while the rest of you is unremarkable [17][18].
– Peripheral hypermobility: concentrated in the hands or the feet, which is a bit of a problem for a score with four of its nine points in the fingers and thumbs [17].
– Historical hypermobility: you had it, it’s genuinely reduced over the years, and the person measuring you today will find very little [17][18].
All four of those can sit alongside significant symptoms, and three of them will produce a low or negative Beighton score right now [17][18], which is written into the framework itself rather than being some loophole somebody found in it.
Did the framework fix the confusion. Partly. hEDS and HSD still overlap a great deal once real people are sat in front of you, and several groups have reported substantial disagreement between what the formal criteria say and the labels people actually walk around with [15][19][20]. Ask a room of people with one of these diagnoses how they got it and you’ll get a room of different stories, and that’s not because anybody is being careless, it’s because the criteria are drawing lines across a continuous trait and lines like that are always going to be arguable [21].
What that means for reading your own paperwork is fairly practical. HSD sits in a different box on the same spectrum rather than being some lesser version of hEDS, and the pain, fatigue and instability in the HSD box can be every bit as bad. Being sorted into one rather than the other says a great deal more about which features happened to be looked for than it does about how unwell you are.
The Score That Counts as Positive Keeps Moving
The threshold you were measured against isn’t fixed, it has never been fully agreed, and the numbers that have been in common use for years are being revised upwards for some groups and downwards for others. Which tends to annoy people, and fairly so.
In children, the older thresholds of 4 or more, or 5 or more, look much too permissive. Pooled data across a very large number of children supports a working cutoff of at least 6 out of 9, and possibly 7 in some groups of girls [22]. That’s a big shift, as a child scored as hypermobile at 5 under the old thinking isn’t hypermobile under the new one, and nothing whatsoever about the child changed.
You can see why the old thresholds were too loose by looking at what ordinary school children score. In one group of Dutch children aged 6 to 12, more than a third scored above 5 out of 9 [23]. More than a third. If a third of a school corridor clears your threshold then the threshold isn’t identifying anything unusual, it’s identifying childhood. Do note that those children weren’t being assessed for symptoms, so this says nothing whatsoever about how many of them had any problems, which is rather the point: the score on its own doesn’t tell you.
Across adulthood, age matters a great deal. In an Australian sample spanning from small children to people over a hundred, a single cutoff of 4 or more applied to everybody produced very poor sensitivity along with a high burden of false positives, depending on which age group you looked at [24]. One number does not fit a lifespan, as tissue behaves differently at 20 and at 70, and a test with a fixed pass mark will get both ends wrong in opposite directions.
The most recent attempt to fix that proposes adult cutoffs that step down with age: 6 or more between 18 and 25, 5 or more between 26 and 65, and 4 or more above 65 [25]. That’s a data driven proposal rather than settled doctrine, and more information is still needed on ethnicity and on thresholds for particular subgroups [25]. Worth holding lightly, then, but it does show the direction of travel, which is away from one universal pass mark and towards a number that knows how old you are.
What all that means for your own score is a bit uncomfortable. If you were tested at 34 and scored 4, you cleared the old flat cutoff and you don’t clear the newer proposed one. If you were tested at 68 and scored 4, it’s the other way round. Same body, same number, different verdict, and the only thing that actually changed is which paper the person testing you had read most recently.
– If you were scored as a child: a 5 that counted as positive a few years ago may well not count now, as the children’s threshold has moved up to 6 or more [22]. That doesn’t mean anybody lied to you, it means the line moved.
– If you were scored as a younger adult: the newer proposal asks for 6 or more between 18 and 25, which is higher than most people have ever been held to [25].
– If you were scored later in life: a flat cutoff of 4 applied across all ages misses a lot of people, and a lower threshold in older age is the sensible reading of what’s actually been measured [24][25].
– If you were scored once, years ago: that number describes that day. Range changes with age, injury, surgery and guarding, and nothing in the test captures where you started from [17].
Now, there’s a knock on effect for anybody who was assessed as a child, and it cuts both ways. If a child was called hypermobile on a 5 under the old thresholds, that label was generous by current standards [22], and given how many ordinary school children clear 5 anyway, a childhood label on its own carries a good deal less information than it appears to [23]. But the reverse is just as true, as a child who scored 5 and was told that meant nothing, in a setting still using a higher bar, has been measured against a threshold that was never about symptoms in the first place. Neither of those children got a wrong number, they got a number read against a moving line, and the line moved without anybody telling them.
And a threshold is somebody’s choice rather than a discovery. Move it up and you catch fewer people who are fine and miss more who aren’t, move it down and you do the reverse, and there’s no setting that gets both right, which is why the number keeps moving and why arguing about a single point either side of it is usually a waste of an appointment.
How Much Hypermobility Varies Between Populations
The score was built on populations, and it shows. Average joint range differs by sex and it differs between populations, so a single cutoff applied to everybody is always going to be measuring different things in different groups.
Women tend to score higher than men, and in that Australian sample, so did people who weren’t of white European background [24]. That pattern turns up repeatedly, and in different directions depending on where anybody looked: high rates of joint laxity in West African groups [26], higher mobility reported in African populations [13], and variation between Malay, Indian, Chinese, white European and African American groups depending on the sample and the cutoff being used [27][28].
So the demographic bias isn’t a suspicion, it’s baked in. Apply one pass mark across unlike populations and you’ll over identify some groups and under identify others, and you’ll do it silently, because the number looks exactly the same on every form.
It’s worth being careful about how far to take that though. These aren’t uniform, tidy effects with a clean ranking of populations, and the sample and the cutoff change the answer, which is precisely why more data on ethnicity is needed before anybody writes subgroup specific numbers down [25]. What you can say with a straight face is narrower and rather more useful: your score has to be read against people like you, and most of the time it isn’t.
That has a practical edge for anybody who’s ever been told their score is “normal for their background”. Normal for a population is a statement about a population. It tells you roughly where you sit on a curve, and it tells you nothing at all about whether your shoulder keeps coming out of its socket.
Five Movements, and the Joints They Miss
The literature is unusually consistent on this one, which happens rarely enough to be worth mentioning. The Beighton score samples a narrow slice of the body, it’s weighted towards the upper limb, and it disregards a lot of major joints, which is why it can’t directly identify generalised hypermobility on its own [5][29][30].
The joints that keep coming up as missing are the ones people actually complain about, so shoulders, hips, ankles and the cervical spine. Those regions appear over and over in critiques of the score and in the design of the tools built to sit alongside it, and if you’ve spent years with a shoulder that wanders and a neck that can’t hold a position, you’ve probably already worked out that nobody measured either of them [10][20][9].
This isn’t a theoretical complaint. In children who met the Beighton definition of hypermobility, only about a quarter turned out to have the spine, the upper limbs and the lower limbs all involved together, and scores of 6 or 8 often reflected arms and legs with no axial involvement at all [31]. Which makes the word generalised do an awful lot of heavy lifting, as a child can clear the threshold for generalised joint hypermobility while a whole region of their body is entirely ordinary.
It runs the other way too. Among young university students, localised hypermobility was more common than the generalised sort, and a positive fifth finger was particularly frequent [32]. Hypermobile little fingers are common, and two of the nine points sit on them with another two on the thumbs, so a substantial chunk of the total can come from a few small joints in the hands while nothing else about you is especially mobile.
Sit those two findings side by side and the score’s real behaviour comes into focus. It can hand you a high number off the back of a few small joints, and it can hand you a low number while the shoulder, hip and neck you actually came in about go completely unexamined. Both of those are wrong in ways that feel very personal when it happens to you.
The lower limb is its own particular blind spot. Foot and ankle mechanics in hypermobile children have been looked at directly, and the sorts of things that turn up there are not things a fifth finger and a hamstring test are ever going to reveal [9]. Anyone who has watched a hypermobile child’s arches collapse while they scored a tidy 4 already knows this.
Which is why other tools exist. The upper limb assessment tool measures several upper limb joints rather than two and a half, and it holds up well between different examiners in adults [33]. Different tests of laxity don’t always agree with each other either, which is its own quiet indictment of treating any single one of them as the answer [30].
Two honest caveats about all of them though. The evidence comparing these tools against each other is limited overall, so they’re complements rather than proven replacements [12]. And more measuring is not automatically better measuring, as a longer examination nobody has validated properly is just a longer examination.
Scoring Low and Still Having Every Symptom
This is the question most people arrive with, and the straight answer is that a low current Beighton score does not rule out a hypermobility disorder.
Three of the four patterns named earlier will produce a low score today. Localised hypermobility, peripheral hypermobility and historical hypermobility are all recognised, all distinct, and all perfectly capable of sitting alongside significant symptoms while the number you’d score this afternoon is low or plainly negative [17][18].
The historical one catches the most people, so it’s worth spelling out. If you could put your palms flat on the floor at twelve, drop into the splits without warming up, and pop your shoulder out to entertain your mates, and you now can’t get past your knees, your tissue didn’t stop being what it is. Your available range changed. The score reads range, so the score changed along with it, and if nobody asks about the twelve year old version of you then that entire history is invisible to the assessment.
There’s reasonable support for taking the other two seriously as well. People with localised and historical hypermobility report symptoms and coexisting conditions that look a lot like what gets reported in hEDS and HSD, which is a strong hint that where your laxity sits, and whether it’s still there to be measured, is a different question from how unwell you are [34]. Take that as self reported, because it is, and overlapping self reported symptoms are a reason to stop dismissing people rather than proof that the underlying condition is identical.
Underneath all of it is the distinction the score simply can’t make. Joint hypermobility is a trait, whereas HSD and hEDS are disorders, and getting to either of them requires symptoms, features across more than one system, exclusion of other explanations, and clinical judgement [14][15][16]. A trait measure cannot diagnose a disorder however carefully you measure the trait, and no amount of precision about nine joints is ever going to change that.
So if you’ve been told you can’t have a hypermobility disorder because you scored 3, what you’ve actually been told is that five movements came out within range on a Tuesday, which is a piece of information rather than a conclusion, and it certainly isn’t an assessment.
The Five Part Questionnaire, and Why History Matters
Given all of that, the obvious fix is to ask about the past rather than only measuring the present, and there’s a short questionnaire that does exactly that. It was built as an adjunct, a self reported screen designed to sit alongside the physical examination rather than to replace it, and it asks about the things you used to be able to do [2].
The five questions cover whether you can now, or could ever, put your hands flat on the floor with your knees straight, whether you can now or could ever bend your thumb to touch your forearm, whether you entertained anybody as a child by contorting your body or doing the splits, whether a shoulder or a kneecap has dislocated on more than one occasion, and whether you’d describe yourself as double jointed [2]. Two or more yes answers points towards generalised joint hypermobility [2].
Look at what those five questions are actually doing, as it’s cleverer than it looks. Two of them are the physical test asked in the past tense, so the palms on the floor and the thumb to forearm from the nine points turn up again with “or could you ever” attached to them. One asks about childhood contortion, which catches people whose flexibility was so unremarkable to them at the time that they’d never once have called it hypermobility. One asks about repeated dislocations, which is a symptom rather than a range of movement and is the only place instability gets a look in at all. And one just asks what you’d call yourself, which sounds deeply unscientific and is doing real work, as being the kid who was “double jointed” is a thing people remember for decades.
It works rather better than a five question form has any right to. In the original groups of adult women it picked up most of those who did have generalised hypermobility and correctly cleared most of those who didn’t, at around 84 and 89 per cent, and when it was checked again it came out at 84 and 80 [2]. That’s decent, and it’s why the thing has stuck around for twenty odd years. It’s especially useful where current range has dropped off, whether from age, injury, surgery or years of guarding, which is exactly the group the physical score handles worst [2].
How well it travels depends on where it goes though. It performed well in Swedish adults who weren’t drawn from a clinical population [35], acceptably in young Chinese adults [36], and it’s been translated and validated in other languages, with the same finding sitting underneath: the measure it keeps getting compared against doesn’t cleanly capture whole body hypermobility in the first place [11]. In early pregnancy it looked rather more uncertain, and that one comes with an interesting wrinkle, as the thing it was tested against was a Beighton score of 5 or more [10]. Judge a questionnaire against a yardstick with known validity problems and a poor result is genuinely ambiguous. Either the questionnaire is wrong, or the yardstick is, and the paperwork can’t tell you which.
Which is a small, specific example of a much bigger problem in this area. Most of the work checking whether a hypermobility measure is any good has to check it against another hypermobility measure, and the best available comparator is the one everybody already agrees is too narrow. It’s a bit like calibrating a set of scales against a set of scales.
Still, as a practical matter, if you’re heading into an appointment and you know your current range isn’t what it was, the five questions are a better description of you than today’s number is. Write your answers down and take them with you, as nobody is going to guess.
When a Screening Tool Becomes a Gate
The Beighton score got a great deal more consequential when it stopped being an epidemiological counting tool and started being the thing that decided who got through a door [5], and that happened without anybody really deciding it.
The people who study the score are now fairly blunt about the limits of that role. It shouldn’t be the principal way of telling localised hypermobility apart from generalised, and it shouldn’t be used on its own to rule generalised joint hypermobility out [5][12].
Used on its own to rule out is, of course, precisely what happens. And there’s a worked example of the cost: cases of classical Ehlers-Danlos syndrome, confirmed at the molecular level, that would have been missed if getting any further assessment had depended on a negative Beighton score [5]. Classical EDS is a different condition from hEDS and it has a genetic test behind it, so those are people whose diagnosis was sitting there available and was very nearly withheld by a nine point flexibility check.
The number is also considerably less stable in real use than it looks. Where people already carrying an hEDS diagnosis were reassessed by specialists, the scores originally taken in primary care came out higher than the specialist scores in most of the cases looked at again, generalised hypermobility was confirmed in fewer than half of those already diagnosed, and, tellingly, a higher score didn’t come with more going on in other body systems [19]. Sit with that last part for a second though, as if the score tracked severity or systemic involvement in any useful way then scoring higher would mean being more unwell, and it didn’t.
Take it gently, as that was one specialist service reviewing its own referrals, and the people who end up at a specialist service are not a cross section of anybody. What it does show is that the number can move depending on who’s holding your fingers, which is awkward for a measure being used as a threshold, and it also shows that the score and the systemic picture are pulling in different directions.
The disagreement goes wider than one service. Globally, people with hEDS and HSD describe a long, difficult diagnostic process, a stack of coexisting conditions, and needs that aren’t being met, which is not the profile of a condition anybody is confidently and consistently identifying [20]. And there’s a genuinely philosophical problem sitting in the middle of all of it, which is that any criteria set drawing a line between normal and pathological on a continuous trait will contain contradictions, and the hEDS criteria have been picked apart on exactly those grounds [21].
None of which is an argument for binning the criteria, it’s an argument for not treating a threshold as though it were a fact about your body, rather than a decision somebody made about where to cut a distribution.
What Gatekeeping Actually Costs
None of this would matter very much if a threshold were only an administrative detail, and unfortunately it isn’t, as the evidence on what happens to people is consistent enough to be worth stating plainly.
Long delays before diagnosis, repeated wrong diagnoses along the way, being dismissed, being met with scepticism, low trust in medical care afterwards, and real emotional distress: that pattern turns up again and again in work with people who have hEDS and HSD, through interviews, through national surveys and across different countries [37][38][39][40]. Some of that work names the mechanism more sharply, describing difficult medical encounters as themselves a source of lasting harm [41].
It’s self reported, and it’s reported by people who eventually got a diagnosis and were therefore available to be asked, which is a real limitation and probably means the picture is, if anything, generous. It’s also, if you’re reading this having lived it, unlikely to be news.
What that looks like from the inside is years, not weeks. Repeated wrong diagnoses aren’t a single administrative error, they’re a sequence of appointments each of which ends with an explanation that doesn’t fit, and the reported pattern includes a long list of coexisting conditions being managed separately by people who never once put them together [38][20]. Somewhere in that sequence a lot of people quietly stop bringing things up, which is a perfectly rational response to being disbelieved and a genuine problem for whoever eventually does take a proper history [37][39].
The other half of that finding gets a great deal less airtime. Getting the diagnosis often brings a sense of validation and better coping afterwards [42], which explains why being included or excluded by a threshold lands so much harder than a number ought to. The label isn’t only administrative, as it changes how you’re treated, what help exists, what gets taken seriously next time, and not least whether you’re believed, and that is exactly why drawing the line a single point further along changes a life [21].
There’s a second cost that gets discussed a lot less, and it points the other way. A score handed out too freely creates its own confusion, as a positive number without symptoms, without a history and without a broader examination isn’t a diagnosis either, and the field’s own reviews keep saying the score can misclassify in both directions [5][12]. Over labelling isn’t kind either, as it fills waiting lists with people who’ve been told they have something nobody has actually established, and it makes the whole area considerably easier to dismiss.
So the score carries consequences a screening tool was never designed to carry. Somebody counting hypermobile joints across a population in the early 1970s was not making a decision about whether you’d be believed in 2026, and yet here we are [13].
Scoring Nine and Being Fine
The other direction gets much less attention and deserves some, as plenty of people arrive at a 9 and assume it settles something.
It doesn’t though, as a high score says your joints are mobile in the movements tested, and that’s the whole of what it says. It doesn’t grade severity and it carries no information about symptoms, function or anything systemic [3][7]. Getting to hEDS or HSD needs symptoms, features in more than one system, other explanations ruled out and clinical judgement, and a 9 supplies none of those [14][16].
There’s supporting evidence for that from a fairly unexpected direction. Among eleven year olds in ordinary schools, joint hypermobility wasn’t associated with musculoskeletal pain or with neurodevelopmental problems [29]. One school based group of children of one age, so please don’t stretch it any further than that, and plenty of adults with hypermobility have pain that’s entirely real and entirely explicable. What it does show is that hypermobile and unwell are not the same population, and a flexible nine year old is not, on their own, a forecast.
Genetic testing doesn’t settle it either, which does surprise people. There’s no genetic test for hEDS, and testing’s real role is identifying the other heritable connective tissue conditions that need ruling out, which is part of why diagnosis here stays difficult and judgement heavy [16].
Where a high score is genuinely useful is as a flag for a conversation rather than as the conclusion to one. If you score 9 and you’re well, you’re mobile and well, and that’s a perfectly good place to be. If you score 9 and you’re not well, the 9 is very nearly the least informative thing about you.
What to Do With Your Own Number
The practical bit, then. Your Beighton score is a reasonable description of how mobile a few joints were on the day somebody measured them, and it isn’t a verdict on whether your symptoms are real, nor is it the deciding factor on whether a hypermobility disorder is possible. What goes into a proper assessment is a good deal broader: historical flexibility, the regions the score never looks at, pain, instability, skin and tissue features, autonomic and gut symptoms, and family history [6][8][18].
– Write down what you used to be able to do: the five part questionnaire exists precisely because history carries information today’s range doesn’t [2]. Palms flat as a child, thumb to forearm, contorting for entertainment, a shoulder or kneecap out more than once, ever called double jointed. Two or more of those is meaningful, and it’s the bit most likely to go unrecorded if nobody thinks to ask.
– Name the joints that aren’t in the test: shoulders, hips, ankles and the neck are the usual omissions and they’re also the usual complaints [10][20]. If your problem joint isn’t one of the five being measured, say so out loud, as the score will not say it for you.
– Ask what threshold you’re being measured against, and at what age: the numbers differ by age and the proposals have moved, so 4 or more at 30 is a different claim from 4 or more at 70 [25][24]. It’s a fair question, and the answer tells you something rather useful about whoever is answering it.
– Bring the rest of the body with you: features across more than one system are part of how HSD and hEDS get recognised, so the gut, the skin, the blood pressure and the fatigue are not side issues to mention if there’s time at the end [14][16].
– Don’t practise for the test: there’s nothing to be gained from forcing an end range you don’t normally have, and shoving a joint past where it wants to go for the sake of a point is a bad trade. It’s supposed to be a passive test anyway [3].
– Ask for the history to go in the notes: a low score with an unrecorded history reads, to the next person who opens the file, as a low score and nothing else [17][18].
What a Proper Assessment Looks Like
For anybody assessing hypermobility rather than being assessed, the supported position is multitool and multisystem rather than Beighton only triage. In practice that means standardised joint testing combined with clinical judgement, a symptom history, a broader joint examination than the nine points allow, the five part questionnaire where current range has dropped off, and active exclusion of overlapping connective tissue conditions and other causes [5][43][44][16]. Primary care is where most of this lands first, and it’s where the practical guidance is aimed, which matters a great deal given how few people ever reach a specialist service [18][44].
Multisystem is the word doing the work there. hEDS and HSD are conditions that turn up across several systems at once rather than joint conditions with a few extras attached, so an assessment that only looks at joints will only ever find joints [14][16]. That’s why the guidance aimed at primary care covers the symptom history and the broader picture rather than just the scoring [18][44], and it’s why exclusion matters as much as inclusion: ruling out the other heritable connective tissue conditions is part of the work rather than an optional extra, and genetic testing earns its place there rather than as a test for hEDS itself [16].
The measurement properties of the various ways of classifying generalised hypermobility have been gone through carefully, and the honest summary is that no single method has everything you’d want from it [43]. Which is a duller conclusion than “throw out the Beighton score” and a much more useful one. You combine things, you write down the history, you look at the regions the standard tests skip, and you accept that judgement is part of it rather than a failure of rigour.
What that means if you’re the one being assessed is that a ten minute appointment ending in a number is the start of an assessment rather than the whole of one, and it’s perfectly reasonable to say so.
And on the technology, which comes up a lot now. Self reported versions of the score, including line drawing formats, can perform well in the right setting [7], and video based systems using pose tracking are being built to help triage people who’ve already been referred [6]. Both are being developed as screening and decision support rather than as stand alone diagnostics, and it’s worth being clear eyed about that, as automating a narrow measure very accurately produces an extremely accurate narrow measure.
What Nobody Knows Yet
A few things genuinely aren’t settled, and saying so is a good deal more useful than pretending otherwise.
Nobody has established which combination of tools does the job better than the Beighton score alone. The alternatives are reasonable and some have decent reliability behind them, but the head to head evidence comparing them is limited, so they sit as complements rather than proven replacements [12][33].
Nobody has worked out the thresholds for particular populations either. The direction is clearly towards age aware cutoffs, and more data on ethnicity and on specific subgroups is needed before anybody writes those numbers down [25]. Until that work exists, reading your score against a population you’re not from remains the default, which is not great.
Nobody has settled the procedure, come to that. Active or passive, goniometer or eyeball, which starting position: those are open questions about a test that’s been in use for fifty years, and they’re open in the sense that different people do it differently rather than in the sense that anybody is arguing about it [3][8].
And there’s no biomarker. No blood test, no scan, nothing that identifies clinically meaningful hypermobility without somebody making a judgement call, which is why hEDS remains a clinical diagnosis with all of the inconsistency that brings [16][15].
What has shifted though, is the thinking. The movement in the literature is away from Beighton only reasoning and towards assessment that takes account of age, takes a history, and looks at more than one body system [5][43][14]. That’s a better way to be looked at and it’s the standard worth asking for.
The score itself is fine at the job it was built for. It’s reproducible, it’s quick, it gives everybody a shared way of describing joint laxity, and it deserves to stay in use for that [3][7]. What it can’t do is map your whole body, grade your symptoms, or stand as the only gate into a diagnosis, and just about every bit of the evidence on it now says so [5][12].
The Fibro Guy

References
[1] Sirajudeen, M.S., Waly, M., Alqahtani, M., Alzhrani, M., Aldhafiri, F., Muthusamy, H. et al. (2020) ‘Generalized joint hypermobility among school-aged children in Majmaah region, Saudi Arabia’, PeerJ. https://doi.org/10.7717/peerj.9682
[2] Hakim, A. and Grahame, R. (2003) ‘A SIMPLE QUESTIONNAIRE TO DETECT HYPERMOBILITY: AN ADJUNCT TO THE ASSESSMENT OF PATIENTS WITH DIFFUSE MUSCULOSKELETAL PAIN’, International Journal of Clinical Practice. https://doi.org/10.1111/j.1742-1241.2003.tb10455.x
[3] Bockhorn, L.N., Vera, A.M., Dong, D., Delgado, D.A., Varner, K.E. and Harris, J.D. (2021) ‘Interrater and Intrarater Reliability of the Beighton Score: A Systematic Review’, Orthopaedic Journal of Sports Medicine. https://doi.org/10.1177/2325967120968099
Read More[4] Schlager, A., Ahlqvist, K., Rasmussen-Barr, E., Bjelland, E.K., Pingel, R., Olsson, C. et al. (2018) ‘Inter- and intra-rater reliability for measurement of range of motion in joints included in three hypermobility assessment methods’, BMC Musculoskeletal Disorders. https://doi.org/10.1186/s12891-018-2290-5
[5] Malek, S., Reinhold, E.J. and Pearce, G.S. (2021) ‘The Beighton Score as a measure of generalised joint hypermobility’, Rheumatology International. https://doi.org/10.1007/s00296-021-04832-4
[6] Sabo, A., Taati, B., Deshpande, A. and Mittal, N. (2026) ‘Development and evaluation of a vision pose-tracking based Beighton score tool for generalized joint hypermobility in individuals with suspected Ehlers-Danlos syndromes’, BioMedical Engineering OnLine. https://doi.org/10.1186/s12938-026-01527-4
[7] Cooper, D.J., Scammell, B.E., Batt, M.E. and Palmer, D. (2018) ‘Development and validation of self-reported line drawings of the modified Beighton score for the assessment of generalised joint hypermobility’, BMC Medical Research Methodology. https://doi.org/10.1186/s12874-017-0464-8
[8] Mittal, N., Sabo, A., Deshpande, A., Clarke, H. and Taati, B. (2022) ‘Feasibility of video-based joint hypermobility assessment in individuals with suspected Ehlers-Danlos syndromes/generalised hypermobility spectrum disorders: a single-site observational study protocol’, BMJ Open. https://doi.org/10.1136/bmjopen-2022-068098
[9] Maarj, M., Pacey, V., Tofts, L., Fellas, A., Clapham, M. and Coda, A. (2025) ‘Lower Limb Biomechanical Observations in Hypermobile Children: An Exploratory Case—Control Study’, International Journal of Environmental Research and Public Health. https://doi.org/10.3390/ijerph22121776
[10] Schlager, A., Ahlqvist, K., Pingel, R., Nilsson-Wikmar, L., Olsson, C.B. and Kristiansson, P. (2020) ‘Validity of the self-reported five-part questionnaire as an assessment of generalized joint hypermobility in early pregnancy’, BMC Musculoskeletal Disorders. https://doi.org/10.1186/s12891-020-03524-7
[11] Moraes, D.A.D., Baptista, C.A., Crippa, J.A.S. and Louzada-Junior, P. (2011) ‘Tradução e validação do The five part questionnaire for identifying hypermobility para a língua portuguesa do Brasil’, Revista Brasileira de Reumatologia. https://doi.org/10.1590/s0482-50042011000100005
[12] Alexander, M. (2022) ‘A Systematised Review of the Beighton Score Compared with Other Commonly Used Measurement Tools for Assessment and Identification of Generalised Joint Hypermobility (GJH)’. https://doi.org/10.1101/2022.04.25.22274226
[13] Beighton, P., Solomon, L. and Soskolne, C.L. (1973) ‘Articular mobility in an African population.’, Annals of the Rheumatic Diseases. https://doi.org/10.1136/ard.32.5.413
[14] Morlino, S. and Castori, M. (2023) ‘Placing joint hypermobility in context: traits, disorders and syndromes’, British Medical Bulletin. https://doi.org/10.1093/bmb/ldad013
[15] Jones, E. and Carrieri, D. (2025) ‘Understanding the issues of hypermobility spectrum disorders and hypermobile Ehlers–Danlos syndrome in primary care: a qualitative integrative review’, Disability and Rehabilitation. https://doi.org/10.1080/09638288.2025.2517246
[16] Forghani, I., See, J. and McGonigle, W.C. (2025) ‘Hypermobile Ehlers–Danlos Syndrome: Diagnostic Challenges and the Role of Genetic Testing’, Genes. https://doi.org/10.3390/genes16050530
[17] Castori, M., Tinkle, B., Levy, H., Grahame, R., Malfait, F. and Hakim, A. (2017) ‘A framework for the classification of joint hypermobility and related conditions’, American Journal of Medical Genetics Part C: Seminars in Medical Genetics. https://doi.org/10.1002/ajmg.c.31539
[18] Atwell, K., Michael, W., Dubey, J., James, S., Martonffy, A., Anderson, S. et al. (2021) ‘Diagnosis and Management of Hypermobility Spectrum Disorders in Primary Care’, The Journal of the American Board of Family Medicine. https://doi.org/10.3122/jabfm.2021.04.200374
[19] McGillis, L., Mittal, N., Santa Mina, D., So, J., Soowamber, M., Weinrib, A. et al. (2019) ‘Utilization of the 2017 diagnostic criteria for hEDS by the Toronto GoodHope Ehlers–Danlos syndrome clinic: A retrospective review’, American Journal of Medical Genetics Part A. https://doi.org/10.1002/ajmg.a.61459
[20] Daylor, V., Griggs, M., Weintraub, A., Byrd, R., Petrucci, T., Huff, M. et al. (2025) ‘Defining the Chronic Complexities of hEDS and HSD: A Global Survey of Diagnostic Challenges, Life-Long Comorbidities, and Unmet Needs’, Journal of Clinical Medicine. https://doi.org/10.3390/jcm14165636
[21] Rosàs Tosas, M. (2025) ‘The Contradictions in the Criteria for Diagnosing Hypermobile Ehlers–Danlos Syndrome as Reflecting Some of the Philosophical Debates about the Threshold between the Normal and the Pathological’, The Journal of Medicine and Philosophy: A Forum for Bioethics and Philosophy of Medicine. https://doi.org/10.1093/jmp/jhaf004
[22] Williams, C.M., Welch, J.J., Scheper, M., Tofts, L. and Pacey, V. (2024) ‘Variability of joint hypermobility in children: a meta-analytic approach to set cut-off scores’, European Journal of Pediatrics. https://doi.org/10.1007/s00431-024-05621-4
[23] Smits-Engelsman, B., Klerks, M. and Kirby, A. (2011) ‘Beighton Score: A Valid Measure for Generalized Hypermobility in Children’, The Journal of Pediatrics. https://doi.org/10.1016/j.jpeds.2010.07.021
[24] Singh, H., McKay, M., Baldwin, J., Nicholson, L., Chan, C., Burns, J. et al. (2017) ‘Beighton scores and cut-offs across the lifespan: cross-sectional study of an Australian population’, Rheumatology. https://doi.org/10.1093/rheumatology/kex043
[25] Tofts, L., Pacey, V., Williams, C.M., Welch, J.J., Shannon, B. and Chan, C. (2026) ‘Generalized Joint Hypermobility in Adults: A Systematic Review With Meta‐Analysis to Identify Data‐Driven Cut‐offs Using the Beighton Score’, Arthritis Care & Research. https://doi.org/10.1002/acr.70017
[26] BIRRELL, F.N., ADEBAJO, A.O., HAZLEMAN, B.L. and SILMAN, A.J. (1994) ‘HIGH PREVALENCE OF JOINT LAXITY IN WEST AFRICANS’, Rheumatology. https://doi.org/10.1093/rheumatology/33.1.56
[27] Seow, C., Chow, P. and Khong, K. (1999) ‘A Study of Joint Mobility in a Normal Population’, Annals of the Academy of Medicine Singapore. https://doi.org/10.47102/annals-acadmedsg.v28n2p231
[28] Flowers, P.P.E., Cleveland, R.J., Schwartz, T.A., Nelson, A.E., Kraus, V.B., Hillstrom, H.J. et al. (2018) ‘Association between general joint hypermobility and knee, hip, and lumbar spine osteoarthritis by race: a cross-sectional study’, Arthritis Research & Therapy. https://doi.org/10.1186/s13075-018-1570-7
[29] Glans, M.R., Aziz, A., Kindgren, E., Knez, R., Landgren, M. and Landgren, V. (2025) ‘No association between joint hypermobility, musculoskeletal pain and neurodevelopmental problems in a school-based sample of 11-year-old children’, BJPsych Open. https://doi.org/10.1192/bjo.2025.10881
[30] Brzozowska, E.K. and Sajewicz, E. (2024) ‘Application of non-parametric correlations to compare the compliance of Beighton and Sachse tests in the assessment of hypermobility based on research of the fitness instructors group’, Journal of Bodywork and Movement Therapies. https://doi.org/10.1016/j.jbmt.2023.11.017
[31] Lamari, M.M., Lamari, N.M., de Medeiros, M.P., Giacomini, M.G., Santos, A.B., de Araújo Filho, G.M. et al. (2024) ‘Generalized Joint Hypermobility: A Statistical Analysis Identifies Non-Axial Involvement in Most Cases’, Children. https://doi.org/10.3390/children11030344
[32] Antonio, D.H. and Magalhaes, C.S. (2018) ‘Survey on joint hypermobility in university students aged 18–25 years old’, Advances in Rheumatology. https://doi.org/10.1186/s42358-018-0008-x
[33] Nicholson, L.L. and Chan, C. (2018) ‘The Upper Limb Hypermobility Assessment Tool: A novel validated measure of adult joint mobility’, Musculoskeletal Science and Practice. https://doi.org/10.1016/j.msksp.2018.02.006
[34] Fairweather, D., Bruno, K.A., Darakjian, A.A., Wilson, F.C., Fliess, J.J., Murphy, E.F. et al. (2025) ‘Localized and historical hypermobile spectrum disorders share self-reported symptoms and comorbidities with hEDS and HSD’, Frontiers in Medicine. https://doi.org/10.3389/fmed.2025.1594796
[35] Glans, M., Humble, M.B., Elwin, M. and Bejerot, S. (2020) ‘Self-rated joint hypermobility: the five-part questionnaire evaluated in a Swedish non-clinical adult population’, BMC Musculoskeletal Disorders. https://doi.org/10.1186/s12891-020-3067-1
[36] Wang, Y., Li, X. and Wang, Y. (2026) ‘Diagnostic validation of the Chinese version of the five-part questionnaire for screening joint hypermobility in young adults’, Scientific Reports. https://doi.org/10.1038/s41598-026-45970-8
[37] Anderson, L.K. and Lane, K.R. (2021) ‘The diagnostic journey in adults with hypermobile Ehlers–Danlos syndrome and hypermobility spectrum disorders’, Journal of the American Association of Nurse Practitioners. https://doi.org/10.1097/jxx.0000000000000672
[38] Halverson, C.M.E., Cao, S., Perkins, S.M. and Francomano, C.A. (2023) ‘Comorbidity, misdiagnoses, and the diagnostic odyssey in patients with hypermobile Ehlers-Danlos syndrome’, Genetics in Medicine Open. https://doi.org/10.1016/j.gimo.2023.100812
[39] Berg, K.M. and Dockrell, D.M. (2026) ‘The lived experience of hypermobile Ehlers–Danlos syndrome and hypermobility spectrum disorders in the United Kingdom: findings from a national cross-sectional survey’, Disability and Rehabilitation. https://doi.org/10.1080/09638288.2026.2646723
[40] Hill, J.C., Cascante, D.C. and Breazeale, J. (2025) ‘Surviving their stripes: the diagnostic odyssey and impact of life with hypermobile Ehlers-Danlos syndrome’, BMJ Connections Clinical Genetics and Genomics. https://doi.org/10.1136/bmjccgg-2025-000044
[41] Halverson, C.M.E., Penwell, H.L. and Francomano, C.A. (2023) ‘Clinician-associated traumatization from difficult medical encounters: Results from a qualitative interview study on the Ehlers-Danlos Syndromes’, SSM – Qualitative Research in Health. https://doi.org/10.1016/j.ssmqr.2023.100237
[42] Wang, Y., Jahani, S., Morel‐Swols, D., Kapely, A., Rosen, A. and Forghani, I. (2024) ‘Patient experiences of receiving a diagnosis of hypermobile Ehlers–Danlos syndrome’, American Journal of Medical Genetics Part A. https://doi.org/10.1002/ajmg.a.63613
[43] Juul‐Kristensen, B., Schmedling, K., Rombaut, L., Lund, H. and Engelbert, R.H.H. (2017) ‘Measurement properties of clinical assessment methods for classifying generalized joint hypermobility—A systematic review’, American Journal of Medical Genetics Part C: Seminars in Medical Genetics. https://doi.org/10.1002/ajmg.c.31540
[44] Pinto-Escalante, D., Rodríguez-Alvarado, M., Contreras-Capetillo, S., González-Herrera, L. and Rubi-Castellanos, R. (2026) ‘A Narrative Review of Hypermobile Ehlers-Danlos Syndrome: Diagnostic Challenges and Opportunities in Primary Care’, Mexican Journal of Medical Research ICSA. https://doi.org/10.29057/mjmr.v14i28.16512


