Operant Conditioning · 操作性条件反射
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| Law of Effect/lɔː ɒv ɪˈfekt/ | 效果律 | xiào guǒ lǜ |
| Reinforcement/ˌriːɪnˈfɔːsmənt/ | 强化 | qiáng huà |
| punishment/ˈpʌnɪʃmənt/ | 惩罚 | chéng fá |
| Positive/ˈpɒzɪtɪv/ | 正 | zhèng |
| negative/ˈneɡətɪv/ | 负 | fù |
| Shaping/ˈʃeɪpɪŋ/ | 塑造 | sù zào |
| primary reinforcer/ˈpraɪməri ˌriːɪnˈfɔːsə/ | 原级强化物 | yuán jí qiáng huà wù |
| secondary reinforcer/ˈsekəndəri ˌriːɪnˈfɔːsə/ | 次级强化物 | cì jí qiáng huà wù |
| successive approximations/səkˈsesɪv əˌprɒksɪˈmeɪʃnz/ | 逐步接近 | zhú bù jiē jìn |
| Continuous reinforcement/kənˈtɪnjuːəs ˌriːɪnˈfɔːsmənt/ | 连续强化 | lián xù qiáng huà |
| ratio/ˈreɪʃɪəʊ/ | 比率 | bǐ lǜ |
| interval/ˈɪntəvl/ | 时距 | shí jù |
| Variable ratio/ˈveərɪəbl ˈreɪʃɪəʊ/ | 可变比率 | kě biàn bǐ lǜ |
| Superstitious behaviour/ˌsuːpəˈstɪʃəs bɪˈheɪvjə/ | 迷信行为 | mí xìn xíng wéi |
| Learned helplessness/lɜːnd ˈhelpləsnəs/ | 习得性无助 | xí dé xìng wú zhù |
Behaviour that pays is behaviour that repeats
- Classical conditioning links two stimuli. Operant conditioning links a behaviour to its consequence.
- The Law of Effect 效果律 states that behaviour followed by a good outcome is more likely to recur.
- The learner is now doing something, not just receiving something.
有回报的行为就是会重复的行为
- 经典条件作用把两个刺激联系起来。操作性条件作用把行为与其结果联系起来。
- 效果律(Law of Effect)指出,后接良好结果的行为更可能再次出现。
- 学习者此时是在做某事,而不只是在接受某事。
Four cells, two questions
- Reinforcement 强化 increases a behaviour; punishment 惩罚 decreases it.
- Positive 正 means something is added; negative 负 means something is removed.
- So negative reinforcement still increases behaviour — by removing something unpleasant.
四个格子,两个问题
- 强化(reinforcement)使行为增加;惩罚(punishment)使行为减少。
- 正(positive)意味着增加某物;负(negative)意味着移除某物。
- 因此负强化仍然使行为增加——通过移除某种不愉快的东西。
Which cell of the grid? · 对应网格中的哪一格?
Sort each consequence by whether it adds or removes, and increases or decreases. · 根据后果是增加/减少以及添加/移除对其进行排序。
Taking a painkiller to remove a headache, and doing it again next time, is an example of... · 服用止痛药去除头痛,并在下次继续这样做,是……的例子
Something unpleasant was removed and the behaviour increased — negative reinforcement. · 某不愉快刺激被移除且行为增加——这是负强化。
Primary, secondary, and shaping
- A primary reinforcer 原级强化物 satisfies a biological need; a secondary reinforcer 次级强化物 is learned, like money.
- Shaping 塑造 builds a complex behaviour by reinforcing successive approximations 逐步接近.
- Shaping is how a behaviour that never occurs spontaneously can still be reinforced.
一级、二级与塑造
- 原级强化物(primary reinforcer)满足生物性需要;次级强化物(secondary reinforcer)是习得的,例如金钱。
- 塑造(shaping)通过强化逐步接近(successive approximations)来建立复杂行为。
- 塑造正是使一个从不自发出现的行为仍能被强化的办法。
Reinforcing successive approximations to build a complex behaviour is called . · 通过强化逐步逼近来构建复杂行为的过程称为。
Shaping is how a behaviour that never occurs spontaneously can be reinforced into existence. · 塑造使得原本不会自发出现的行为也能被强化而得以产生。
Select all · 所有 that are true of a secondary reinforcer. · 选择所有关于次级强化的正确描述。
A primary · 第一产业 (primary) reinforcer satisfies a biological need directly; a secondary one is learned. · 初级强化物直接满足生理需求;次级强化物则是习得的。
Negative reinforcement is not punishment. Negative means something is removed; reinforcement means the behaviour increases. Taking an aspirin to remove a headache is negative reinforcement — and you are more likely to do it again.
负强化不是惩罚。"负"意味着移除某物;"强化"意味着行为增加。吃一片阿司匹林以消除头痛就是负强化——而你此后更可能再这样做。
Which schedule produces the highest and most extinction-resistant response rate? · 哪种强化程序能产生最高且最抗消退的反应率?
Variable ratio — an unpredictable number of responses per reward. · 可变比率——每次奖励所需的反应次数不可预测。
Schedules of reinforcement
- Continuous reinforcement 连续强化 rewards every response — fast learning, fast extinction.
- Partial schedules are fixed or variable, by ratio 比率 (number of responses) or interval 时距 (time).
- Variable ratio 可变比率 produces the highest, most extinction-resistant rate of responding.
强化程式
- 连续强化(continuous reinforcement)奖励每一次反应——学得快,消退也快。
- 部分强化程式分为固定与可变,并按比率(ratio,反应次数)或时距(interval,时间)划分。
- 可变比率(variable ratio)产生最高、最难消退的反应率。
Superstitious behaviour occurs when a consequence reinforces a behaviour that did not cause it. · 当后果强化了一个并未导致该后果的行为时,就会发生迷信行为。
The learner tracks what followed, not what actually caused the outcome. · 学习者关注的是随后发生了什么,而非真正导致结果的原因。
When consequences mislead
- Superstitious behaviour 迷信行为 occurs when a consequence reinforces an unrelated behaviour.
- Learned helplessness 习得性无助 occurs when repeated inescapable outcomes teach that nothing one does matters.
- Both show that the learner tracks what followed, not what actually caused it.
当结果具有误导性时
- 迷信行为(superstitious behaviour)出现在某一结果强化了与之无关的行为时。
- 习得性无助(learned helplessness)出现在反复无法逃避的结果教会人"做什么都没用"时。
- 两者都表明学习者追踪的是随后发生了什么,而不是实际上是什么造成了它。
A slot machine pays out after an unpredictable number of pulls — a variable-ratio schedule. That is exactly the schedule that produces the highest response rate and the greatest resistance to extinction, which is why the design is not accidental.
老虎机在拉动次数不可预测时给出回报——这是可变比率程式。而这恰恰是产生最高反应率、最难消退的那种程式,因此这样的设计并非偶然。
Operant conditioning links behaviour to consequence via the Law of Effect. Reinforcement increases and punishment decreases; positive adds and negative removes. Shaping builds new behaviour by successive approximations, and variable-ratio schedules produce the most persistent responding.
操作性条件作用通过效果律把行为与结果联系起来。强化使行为增加,惩罚使其减少;正是增加某物,负是移除某物。塑造通过逐步接近建立新行为,而可变比率程式产生最持久的反应。